Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “storage throughput”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Combined Experimental and Computational Efforts to Establish Ion Mobility, Solubility and Stability of Functional Liquids for Electrochemical Energy Storage

This work provides a computation-driven investigation of the stability of organic electrolytes for lithium-air batteries. Electrolyte instability is currently a key challenge that limits practical use of aprotic Li-air batteries, and the chemical processes that cause this instability are often kinetically-driven. Computational screening for kinetic stability involves the determination of reaction barriers for the numerous potential reaction mechanisms, barriers that are challenging to calculate due to the difficulty of locating transition state structures. Here we screen a broad set of substituted electrolytes for susceptibility to nucleophilic attack by superoxide. We find that carbonates are not typically expected to be stable and that sulfones are generally stable, validating literature trends. We study the effects of chemical functionalization with electron-donating and withdrawing groups and their interplay with steric factors, identifying functional groups and other chemical modifications that increase stability in these groups. User-input driven transition state identification is used for these initial calculations, and an automated computational pipeline is subsequently presented and validated as a means to perform further high-throughput searches across mechanisms and chemistries. The pipeline integrates cheminformatics-based reaction encoding, relaxed potential energy scans, and nudged elastic band calculations for an end-to-end approach to barrier calculations. We review this automated search approach and its current limitations, and discuss challenges and further work.

25 ENERGY STORAGE↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

Sparse matrix‐vector and matrix‐multivector products for the truncated SVD on graphics processors

Summary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm.

Aliaga, José I.↗

An integrated high-throughput robotic platform and active learning approach for accelerated discovery of optimal electrolyte formulations

Solubility of redox-active molecules is an important determining factor of the energy density in redox flow batteries. However, the advancement of electrolyte materials discovery has been constrained by the absence of extensive experimental solubility datasets, which are crucial for leveraging data-driven methodologies. In this study, we design and investigate a highly automated workflow that synergizes a high-throughput experimentation platform with a state-of-the-art active learning algorithm to significantly enhance the solubility of redox-active molecules in organic solvents. Our platform identifies multiple solvents that achieve a remarkable solubility threshold exceeding 6.20 M for the archetype redox-active molecule, 2,1,3-benzothiadiazole, from a comprehensive library of more than 2000 potential solvents. Significantly, our integrated strategy necessitates solubility assessments for fewer than 10% of these candidates, underscoring the efficiency of our approach. Our results also show that binary solvent mixtures, particularly those incorporating 1,4-dioxane, are instrumental in boosting the solubility of 2,1,3-benzothiadiazole. Beyond designing an efficient workflow for developing high-performance redox flow batteries, our machine learning-guided high-throughput robotic platform presents a robust and general approach for expedited discovery of functional materials.

25 ENERGY STORAGE↗

PINN surrogate of Li-ion battery models for parameter inference, Part I: Implementation and multi-fidelity hierarchies for the single-particle model

To plan and optimize energy storage demands that account for Li-ion battery aging dynamics, techniques need to be developed to diagnose battery internal states accurately and rapidly. Here, this study seeks to reduce the computational resources needed to determine a battery's internal states by replacing physics-based Li-ion battery models - such as the single-particle model (SPM) and the pseudo-2D (P2D) model - with a physics-informed neural network (PINN) surrogate. The surrogate model makes high-throughput techniques, such as Bayesian calibration, tractable to determine battery internal parameters from voltage responses. This manuscript is the first of a two-part series that introduces PINN surrogates of Li-ion battery models for parameter inference (i.e., state-of-health diagnostics). In this first part, a method is presented for constructing a PINN surrogate of the SPM. A multi-fidelity hierarchical training, where several neural nets are trained with multiple physics-loss fidelities is shown to significantly improve the surrogate accuracy when only training on the governing equation residuals. The implementation is made available in a companion repository (https://github.com/NREL/PINNSTRIPES). The techniques used to develop a PINN surrogate of the SPM are extended in Part II for the PINN surrogate for the P2D battery model, and explore the Bayesian calibration capabilities of both surrogates.

25 ENERGY STORAGE↗

Chemical factors controlling the behaviour of oxide cathodes in batteries

Oxide cathodes enable high-energy lithium-ion and sodium-ion batteries, with their performances fundamentally governed by three interrelated chemical factors: electronic configuration, chemical bonding, and chemical reactivity. Here, we illustrate how these factors dictate the redox energy, structural stability, ionic and electronic transport, and interfacial behavior in both layered oxide and polyanion oxide cathodes. We discuss how crystal-field effects and octahedral-site stabilization energies influence cation migration, and how inductive effects tune bond covalency and operating voltages. We also explain how chemical bonding governs thermal stability, gas evolution, and first-cycle capacity loss, and how alignment of transition-metal redox band with the oxygen 2p band determines electrolyte reactivity. Comparison between lithium and sodium layered oxides further reveals how differences in Li-O and Na-O bond ionicity affect chemical reactivity. Finally, we outline strategies including compositional tuning, surface doping, and electrolyte optimization, and emphasize how high-throughput, data-driven approaches in guiding the design of next-generation oxide cathodes.

25 ENERGY STORAGE↗

Robotic sample changers for macromolecular X-ray crystallography and biological small-angle X-ray scattering at the National Synchrotron Light Source II

In this work, we present two robotic sample changers integrated into the experimental stations for the macromolecular crystallography (MX) beamlines AMX and FMX, and the biological small-angle scattering (bioSAXS) beamline LiX. They enable fully automated unattended data collection and remote access to the beamlines. The system designs incorporate high-throughput, versatility, high-capacity, resource sharing and robustness. All systems are centered around a six-axis industrial robotic arm coupled with a force torque sensor and in-house end effectors (grippers). They have the same software architecture and the facility standard EPICS-based BEAST alarm system. The MX system is compatible with SPINE bases and Unipucks. It comprises a liquid nitrogen dewar holding 384 samples (24 Unipucks) and a stay-cold gripper, and utilizes machine vision software to track the sample during operations and to calculate the final mount position on the goniometer. The bioSAXS system has an in-house engineered sample storage unit that can hold up to 360 samples (20 sample holders) which keeps samples at a user-set temperature (277 K to 300 K). The MX systems were deployed in early 2017 and the bioSAXS system in early 2019.

36 MATERIALS SCIENCE↗

Compact ammonia reforming at low temperature using catalytic membrane reactors

Ammonia is a leading carrier for the storage and transport of renewable hydrogen, but its deployment requires scalable technologies for efficient decomposition and purification. In this work, we report on the efficient delivery of high purity hydrogen from ammonia decomposition using a catalytic membrane reactor (CMR). Improvements to the electroless plating process reduced the Pd membrane thickness by >35%, resulting in commensurate increases in hydrogen permeance without sacrificing selectivity. To increase throughput a commercial Ru/Al 2 O 3 catalyst was added to the lumen, and the CMR could process ammonia flowrates 10–50 times higher than an equivalent packed bed reactor while maintaining the same level of conversion. It is shown that the earth-abundant zeolite clinoptilolite could reduce ammonia impurities in the permeated H 2 to the levels required by PEM fuel cells (<25 ppb). Performance increased significantly across a >500-h durability test due to improvements in membrane permeability. Finally, the results show that CMRs are a viable technology for distributed production of hydrogen from ammonia.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Revisiting Huffman Coding: Toward Extreme Performance on Modern GPU Architectures

Today's high-performance computing (HPC) applications are producing vast volumes of data, which are challenging to store and transfer efficiently during the execution, such that data compression is becoming a critical technique to mitigate the storage burden and data movement cost. Huffman coding is arguably the most efficient Entropy coding algorithm in information theory, such that it could be found as a fundamental step in many modern compression algorithms such as DEFLATE. On the other hand, today's HPC applications are more and more relying on the accelerators such as GPU on supercomputers, while Huffman encoding suffers from low throughput on GPUs, resulting in a significant bottleneck in the entire data processing. In this paper, we propose and implement an efficient Huffman encoding approach based on modern GPU architectures, which addresses two key challenges: (1) how to parallelize the entire Huffman encoding algorithm, including codebook construction, and (2) how to fully utilize the high memory-bandwidth feature of modern GPU architectures. The detailed contribution is fourfold. (1) We develop an efficient parallel codebook construction on GPUs that scales effectively with the number of input symbols. (2) We propose a novel reduction based encoding scheme that can efficiently merge the codewords on GPUs. (3) We optimize the overall GPU performance by leveraging the state-of-the-art CUDA APIs such as Cooperative Groups. (4) We evaluate our Huffman encoder thoroughly using six real-world application datasets on two advanced GPUs and compare with our implemented multi-threaded Huffman encoder. Experiments show that our solution can improve the encoding throughput by up to 5.0x and 6.8x on NVIDIA RTX 5000 and V100, respectively, over the state-of-the-art GPU Huffman encoder, and by up to 3.3x over the multi-thread encoder on two 28-core Xeon Platinum 8280 CPUs.

Tian, Jiannan↗

Revisiting Huffman Coding: Toward Extreme Performance on Modern GPU Architectures

Today's high-performance computing (HPC) applications are producing vast volumes of data, which are challenging to store and transfer efficiently during the execution, such that data compression is becoming a critical technique to mitigate the storage burden and data movement cost. Huffman coding is arguably the most efficient Entropy coding algorithm in information theory, such that it could be found as a fundamental step in many modern compression algorithms such as DEFLATE. On the other hand, today's HPC applications are more and more relying on the accelerators such as GPU on supercomputers, while Huffman encoding suffers from low throughput on GPUs, resulting in a significant bottleneck in the entire data processing. In this paper, we propose and implement an efficient Huffman encoding approach based on modern GPU architectures, which addresses two key challenges: (1) how to parallelize the entire Huffman encoding algorithm, including codebook construction, and (2) how to fully utilize the high memory-bandwidth feature of modern GPU architectures. The detailed contribution is fourfold. (1) We develop an efficient parallel codebook construction on GPUs that scales effectively with the number of input symbols. (2) We propose a novel reduction based encoding scheme that can efficiently merge the codewords on GPUs. (3) We optimize the overall GPU performance by leveraging the state-of-the-art CUDA APIs such as Cooperative Groups. (4) We evaluate our Huffman encoder thoroughly using six real-world application datasets on two advanced GPUs and compare with our implemented multithreaded Huffman encoder. Experiments show that our solution can improve the encoding throughput by up to 5.0× and 6.8× on NVIDIA RTX 5000 and V100, respectively, over the state-of-the-art GPU Huffman encoder, and by up to 3.3× over the multithread encoder on two 28-core Xeon Platinum 8280 CPUs.

Tian, Jiannan↗

Self-Driving Microscopy for AI/ML-Enabled Physics Discovery and Materials Optimization

Materials are the bedrock of economy and foundation for all real-world technologies. The viability of space travel, grid energy storage, solar to fuels conversion, methane removal, and photovoltaic energy solutions hinge on the discovery and optimization of novel materials and rapid scaling toward manufacturing. The last 20 years have seen an exponential growth in the theoretical predictive capability for crystalline materials and small molecules. However, it is only in the last five years that we have seen the rapid expansion of high-throughput synthesis enabled by laboratory robotics and microfluidics, as well as a resurgence of combinatorial synthesis (Abolhasani and Kumacheva 2023; Epps and Abolhasani 2021; Jiang et al. 2022; Rajan 2008; Soldatov et al. 2021; Szymanski et al. 2023). Combinatorial synthesis, microfluidics, and ultimately dip-pen megalibraries have demonstrated the ability to “write” multicomponent nanomaterials at high throughput scale, generating millions of material examples in the 3D, 4D, and 5D composition spaces (Chen et al. 2016, 2019; Jibril et al. 2022).

36 MATERIALS SCIENCE↗

In situ thermal conductivity measurement revealing kinetics of thermochemical reactions

Utilizing thermochemical reactions for thermal energy storage and solar fuel production has been an emerging research topic. Thermal transport properties of the materials are an important parameter that can determine the kinetics and efficiency of thermochemical reactions. With the increasing number of new thermochemical materials (TCMs); however, there is a lack of reliable techniques to monitor the thermal transport property of the materials and their changes as a function of reactions in real time. In this work, we report the in situ monitoring of thermochemical reactions using modulated photothermal radiometry (MPR). The thermal conductivities of two TCMs, namely, calcium hydroxide (Ca(OH) 2 ) and Ba 0.15 Sr 0.85 FeO 3–δ (BSF1585), were measured as a function of temperature and time using the MPR technique. The measured thermal conductivities were correlated to the reaction. The work has two significant contributions to the research communities. First, it provides a non-invasive diagnostic tool for monitoring the thermal transport properties of TCMs that can potentially be a high-throughput measurement technique conducive to optimizing TCMs, reactors, and related thermal systems. Second, for TCMs that show observable changes in thermal transport properties, a correlation between the measured thermal conductivity and the conversion fraction of the reaction can be established for monitoring the reaction kinetics based on thermal characterization.

14 SOLAR ENERGY↗

Maximising the investment returns of a grid‐connected battery considering degradation cost

Energy storage systems (ESSs) are being deployed widely due to numerous benefits including operational flexibility, high ramping capability, and decreasing costs. This study investigates the economic benefits provided by battery ESSs when they are deployed for market‐related applications, considering the battery degradation cost. A comprehensive investment planning framework is presented, which estimates the maximum revenue that the ESS can generate over its lifetime and provides the necessary tools to investors for aiding the decision making process regarding an ESS project. The applications chosen for this study are energy arbitrage and frequency regulation. Lithium‐ion batteries are considered due to their wide popularity arising from high efficiency, high energy density, and declining costs. A new degradation cost model based on energy throughput and cycle count is developed for Lithium‐ion batteries participating in electricity markets. The lifetime revenue of ESS is calculated considering battery degradation and a cost–benefit analysis is performed to provide investors with an estimate of the net present value, return on investment and payback period. The effect of considering the degradation cost on the estimated revenue is also studied. The proposed approach is demonstrated on the IEEE Reliability Test System and historical data from PJM Interconnection.

25 ENERGY STORAGE↗

Self-Forming Thin Interphases and Electrodes Enabling 3-D Structured High Energy Density Batteries

An electrolytically in-situ formed fluoride/lithium based battery has been developed to offer a pathway to scalable reconfigurable solid state batteries of high energy density. Research into the development of novel in-situ formed chemistries encompassing the negative and positive reactive current collectors, and the bi-ion glass conductor along with electrode structure was accomplished with a focused attention on transport. The solid state in-situ batteries were fabricated with a maskless scalable patterning technique to offer a pathway to high throughput, low material loss and fabrication of complex architectures. Such development and integration enabled to achieve the 12 V bipolar batteries at > 1000 Wh/L energy density based on electrode pairs and current collectors.

25 ENERGY STORAGE↗

ThunderSecure: deploying real-time intrusion detection for 100G research networks by leveraging stream-based features and one-class classification network

Nowadays, data generated by large-scale scientific experiments are on the scale of petabytes per month. These data are transferred through dedicated high-bandwidth networks (40/100G) across distributed sites for processing, storage, and analysis. Like general purpose networks, research networks experience intrusions. However, monitoring anomalies in such high-speed network traffics is challenging given current cyber-infrastructure. Moreover, traditional network intrusion detection systems (NIDS) are signature based. However, anomaly patterns are difficult to define and that rulesets are often not updated frequently enough to reflect the changes of attack behaviors. We present ThunderSecure, a high-throughput, unsupervised learning-based intrusions detection system for 100G research networks. ThunderSecure implements an efficient packet processing and detection pipeline using multi-cores and GPUs. It extracts statistical and temporal features from real-time network data streams and feeds them to a one-class anomaly detection network. A baseline of normal distribution will be created based on the training observation. Testing traffic deviated from the learned profile will be marked as anomalies. We trained ThunderSecure on hundreds of billions of science data packets mirrored from two 100G network connections at Fermi National Accelerator Laboratory. The detection performance was evaluated on traffic captured from the same research network days and weeks after the training with different types of attack flows injected. Results show that ThunderSecure can recognize science data traffic captured long after the training and made nearly certain detection on the segment of the streams where anomalous flows were injected.

100G research network↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

High Throughput Solvent-free Manufacturing of Battery Electrodes

The project goal is to develop and demonstrate an advanced solvent-free lithium-ion battery electrode process through proposed Advanced Dry Electrode Process (ADEP) equipment, which is expected to exhibit a better binder fibrillization and high throughput and suitable for high performance electrode manufacturing, and commonize the anode and cathode dry processing equipment, supply chain and operation for lithium-ion battery OEMs for replacing the solvent-based slurry casting. Our proposed approach will facilitate low-cost battery production by addressing the following gaps in present dry electrode processing: • Extend the dry electrode fabrication process to lithium-ion battery anode manufacturing • Increase the active material content for anodes and cathodes • Intensify the process through improved mixing, powder rheology and surface modifications • Enable processing of next-generation electrode materials that are not stable to solvent or ambient air exposure. The project objectives include the development of anode-compatible binder and binder fibrillization promoter for low irreversible capacity loss, low electrode binder content yet higher film mechanical strength, and the optimization of solvent-free anode and cathode process for low cost (>60% electrode cost reduction), high performance (10% increase in energy density without sacrificing cycle life) and high throughput to enable next generation lithium-ion battery electrode production. Solvent-free electrode manufacturing will also enable next-generation cell designs based on prelithiated anodes or solid-state electrolytes.

25 ENERGY STORAGE↗