Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “bottleneck structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Computing Bottleneck Structures at Scale for High-Precision Network Performance Analysis

The Theory of Bottleneck Structures is a recently-developed framework for studying the performance of data networks. It describes how local perturbations in one part of the network propagate and interact with others. This framework is a powerful analytical tool that allows network operators to make accurate predictions about network behavior and thereby optimize performance. Previous work implemented a software package for bottleneck structure analysis, but applied it only to toy examples. In this work, we introduce the first software package capable of scaling bottleneck structure analysis to production-size networks. Here, we benchmark our system using logs from ESnet, the Department of Energy's high-performance data network that connects research institutions in the U.S. Using the previously published tool as a baseline, we demonstrate that our system achieves vastly improved performance, constructing the bottleneck structure graphs in 0.21 s and calculating link derivatives in 0.09 s on average. We also study the asymptotic complexity of our core algorithms, demonstrating good scaling properties and strong agreement with theoretical bounds. These results indicate that our new software package can maintain its fast performance when applied to even larger networks. They also show that our software is efficient enough to analyze rapidly changing networks in real time. Overall, we demonstrate the feasibility of applying bottleneck structure analysis to solve practical problems in large, real-world data networks.

benchmark↗

GradientGraph

Under this SBIR Phase II, Reservoir Labs has developed G2 Analytics, a new technology that allows network operators to analyze bottleneck and flow performance with high precision. G2 delivers a new analytical approach and framework to resolve a variety of key problems found in modern communication networks, including: traffic engineering, routing, flow scheduling, network design, capacity planning, resiliency analysis, network slicing, or service level agreement (SLA) management, among others. G2 leverages the bottleneck structure of congestion-controlled communication networks, a recent mathematical discovery by the Reservoir team [RL19b, RL20a, RL20b, RL21a]. Bottleneck structures reveal how perturbations on flows and links propagate through the network, providing an analytical framework to measure (qualitatively and quantitatively) the ripple effects induced as they traverse the network. Leveraging the mathematics of bottleneck structures, Reservoir Labs is developing the G2 technology to provide network operators with a framework to design, optimize and troubleshoot network performance. This delivery includes the G2 software stack.

Yellamraju, Sruthi↗

Systems and methods for quality of service (QoS) based management of bottlenecks and flows in networks

Techniques based on the Theory of Bottleneck Ordering can reveal the bottleneck structure of a network, and the Theory of Flow ordering can take advantage of the revealed bottleneck structure to manage and configure network flows so as to improve the overall network performance. These two techniques provide insights into the inherent topological properties of a network at least in three areas: (1) identification of the regions of influence of each bottleneck; (2) the order in which bottlenecks (and flows traversing them) may converge to their steady state transmission rates in distributed congestion control algorithms; and (3) the design of optimized traffic engineering policies.

97 MATHEMATICS AND COMPUTING↗

TeraChem: A graphical processing unit-accelerated electronic structure package for large-scale ab initio molecular dynamics

TeraChem was born in 2008 with the goal of providing fast on-the-fly electronic structure calculations to facilitate ab initio molecular dynamics studies of large biochemical systems such as photoswitchable proteins and multichromophoric antenna complexes. Originally developed for videogaming applications, graphics processing units (GPUs) offered a low-cost parallel computer architecture that became more accessible for general-purpose GPU computing with the release of CUDA in 2007. The evaluation of the electron repulsion integrals (ERIs) is a major bottleneck in electronic structure codes and provides an attractive target for acceleration on GPUs. Thus, highly efficient routines for evaluation of and contractions between the ERIs and density matrices were implemented in TeraChem. Here, electronic structure methods were developed and implemented to leverage these integral contraction routines, resulting in the first quantum chemistry package designed from the ground up for GPUs. This GPU acceleration makes TeraChem capable of performing large-scale ground and excited state calculations in the gas and condensed phase. Today, TeraChem's speed forms the basis for a suite of quantum chemistry applications, including optimization and dynamics of proteins, automated and interactive chemical discovery tools, and large-scale nonadiabatic dynamics simulations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Elucidating the Interfacial Barriers in Lanthanide Back-Extraction: From Water to Oil and Back Again

Recovery of critical rare earth elements from complex mixtures has long been realized via solvent extraction, where ions in an aqueous phase are separated into an organic phase using amphiphilic ligands. While a great deal of effort has been placed on understanding this forward reaction, substantial knowledge gaps in the back-extraction process remain. This includes the mechanism of interfacial dissociation and transport back into a highly acidic aqueous phase for further processing. In this work, we connect back-extraction kinetics made in realistic solvent extraction systems to salient interfacial chemistry and structure that represent bottlenecks in the back-extraction of lanthanide ions. We show that the interface between the two liquid phases varies dramatically based on the composition of both phases. Water stretching signals are shown to report on the population of lingering interfacial complexes and are thus used as a reporter of competitive adsorption from excess free ligands in solution for limited interfacial vacancies. We show that excess free ligands, often used to improve forward extractions, set up interfacial blockades inhibiting back-extraction both kinetically and thermodynamically. In conclusion, this insight opens up avenues to tune interfacial properties to facilitate a more dynamic, exchangeable interface to speed up back-extractions while using less energy intensive chemical swings.

Interfaces↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

An Intelligent Distributed Ledger Construction Algorithm for IoT

Blockchain is the next generation of secure data management that creates near-immutable decentralized storage. Secure cryptography created a niche for blockchain to provide alternatives to well-known security compromises. However, design bottlenecks with traditional blockchain data structures scale poorly with increased network usage and are extremely computation-intensive. This made the technology difficult to combine with limited devices, like those in Internet of Things networks. In protocols like IOTA, replacement of blockchain's linked-list queue processing with a lightweight dynamic ledger showed remarkable throughput performance increase. However, current stochastic algorithms for ledger construction suffer distinct trade-offs between efficiency and security. This work proposed a machine-learning approach with a multi-arm bandit that resolved these issues and was designed for auditing on limited devices. This algorithm was tested in a reinforcement-learning environment simulating the IOTA ledger's construction with a decision tree. This study showed through regret analysis and experimentation that this approach was secure against impulse manipulation attacks while remaining energy-efficient. Although the IOTA protocol was a pioneer for lightweight distributed ledgers, it is expected that future blockchain protocols will adopt techniques similar to those presented in this work.

multi-arm bandit↗

Roll-to-roll solvent-free manufactured electrodes for fast-charging batteries

In response to the growing demand for lithium-ion batteries (LIBs), we demonstrate a solvent-free manufacturing technology that can avoid toxic organic solvents and form unique electrode structures to overcome the bottlenecks in low costs and fast charging. The lower tortuosity achieved by the open pores in the dry-printed (DP) electrode allows for a shorter Li+ diffusion pathway, which leads to better rate performance. The DP pouch cells exhibit higher capacity retention of 78% and 69% at 3C and 4C, respectively, compared with 67% and 52% for the slurry cast (SL) cells at the same rates. Moreover, the coating layer on the surface of active materials prevents the excess side reaction between active materials and electrolytes, which prolongs the cycle life of the DP cells. This manufacturing process is a roll-to-roll system with immense potential to be scaled up, providing a more efficient and economical way for battery manufacturing.

36 MATERIALS SCIENCE↗

Dominant Energy Carrier Transitions and Thermal Anisotropy in Epitaxial Iridium Thin Films

High aspect ratio metal nanostructures are commonly found in a broad range of applications such as electronic compute structures and sensing. The self-heating and elevated temperatures in these structures, however, pose a significant bottleneck to both the reliability and clock frequencies of modern electronic devices. Any notable progress in energy efficiency and speed requires fundamental and tunable thermal transport mechanisms in nanostructured metals. Here, in this work, time-domain thermoreflectance is used to expose cross-plane quasi-ballistic transport in epitaxially grown metallic Ir(001) interposed between Al and MgO(001). Thermal conductivities ranges from roughly 65 (96 in-plane) to 119 (122 in-plane) W m -1 K -1 for 25.5–133.0 nm films, respectively. Further, low defects afforded by epitaxial growth are suspected to allow the observation of electron–phonon coupling effects in sub-20 nm metals with traditionally electron-mediated thermal transport. Via combined electro-thermal measurements and phenomenological modeling, the transition is revealed between three modes of cross-plane heat conduction across different thicknesses and an interplay among them: electron dominant, phonon dominant, and electron–phonon energy conversion dominant. The results substantiate unexplored modes of heat transport in nanostructured metals, the insights of which can be used to develop electro-thermal solutions for a host of modern microelectronic devices and sensing structures.

36 MATERIALS SCIENCE↗

Graph-Directed Approach for Downselecting Toxins for Experimental Structure Determination

Conotoxins are short, cysteine-rich peptides of great interest as novel therapeutic leads and of great concern as lethal biological agents due to their high affinity and specificity for various receptors involved in neuromuscular transmission. Currently, of the approximately 6000 known conotoxin sequences, only about 3% have associated structural characterization, which leads to a bottleneck in rapid high-throughput screening (HTS) for identification of potential leads or threats. In this work, we combine a graph-based approach with homology modeling to expand the library of conotoxin structures and to identify those conotoxin sequences that are of the greatest value for experimental structural characterization. The latter would allow for the rapid expansion of the known structural space for generating high quality template-based models. Our approach generalizes to other evolutionarily-related, short, cysteine-rich venoms of interest. Overall, we present and validate an approach for venom structure modeling and experimental guidance and employ it to produce a 290%-larger library of approximate conotoxin structures for HTS. We also provide a set of ranked conotoxin sequences for experimental structure determination to further expand this library.

59 BASIC BIOLOGICAL SCIENCES↗

A goldilocks computational protocol for inhibitor discovery targeting DNA damage responses including replication-repair functions

While many researchers can design knockdown and knockout methodologies to remove a gene product, this is mainly untrue for new chemical inhibitor designs that empower multifunctional DNA Damage Response (DDR) networks. Here, we present a robust Goldilocks (GL) computational discovery protocol to efficiently innovate inhibitor tools and preclinical drug candidates for cellular and structural biologists without requiring extensive virtual screen (VS) and chemical synthesis expertise. By computationally targeting DDR replication and repair proteins, we exemplify the identification of DDR target sites and compounds to probe cancer biology. Our GL pipeline integrates experimental and predicted structures to efficiently discover leads, allowing early-structure and early-testing (ESET) experiments by many laboratories. By employing an efficient VS protocol to examine protein-protein interfaces (PPIs) and allosteric interactions, we identify ligand binding sites beyond active sites, leveraging in silico advances for molecular docking and modeling to screen PPIs and multiple targets. A diverse 3,174 compound ESET library combines Diamond Light Source DSI-poised, Protein Data Bank fragments, and FDA-approved drugs to span relevant chemotypes and facilitate downstream hit evaluation efficiency for academic laboratories. Two VS per library and multiple ranked ligand binding poses enable target testing for several DDR targets. This GL library and protocol can thus strategically probe multiple DDR network targets and identify readily available compounds for early structural and activity testing to overcome bottlenecks that can limit timely breakthrough drug discoveries. By testing accessible compounds to dissect multi-functional DDRs and suggesting inhibitor mechanisms from initial docking, the GL approach may enable more groups to help accelerate discovery, suggest new sites and compounds for challenging targets including emerging biothreats and advance cancer biology for future precision medicine clinical trials.

59 BASIC BIOLOGICAL SCIENCES↗

Neural network acceleration of large-scale structure theory calculations

Here, we make use of neural networks to accelerate the calculation of power spectra required for the analysis of galaxy clustering and weak gravitational lensing data. For modern perturbation theory codes, evaluation time for a single cosmology and redshift can take on the order of two seconds. In combination with the comparable time required to compute linear predictions using a Boltzmann solver, these calculations are the bottleneck for many contemporary large-scale structure analyses. Here, we construct neural network-based surrogate models for Lagrangian perturbation theory (LPT) predictions of matter power spectra, real and redshift space galaxy power spectra, and galaxy-matter cross power spectra that attain ~ 0.1% (at one sigma) accuracy over a broad range of scales in a ωCDM parameter space. The neural network surrogates can be evaluated in approximately one millisecond, a factor of 1000 times faster than the full Boltzmann code and LPT computations. In a simulated full-shape redshift space galaxy power spectrum analysis, we demonstrate that the posteriors obtained using our surrogates are accurate compared to those obtained using the full LPT model. We make our surrogate models public at https://github.com/sfschen/EmulateLSS, so that others may take advantage of the speed gains they provide to enable rapid iteration on analysis settings, something that is essential in complex contemporary large-scale structure analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Vibrational heat-bath configuration interaction with semistochastic perturbation theory using harmonic oscillator or VSCF modals

Vibrational heat-bath configuration interaction (VHCI)—a selected configuration interaction technique for vibrational structure theory—has recently been developed in two independent works [J. H. Fetherolf and T. C. Berkelbach, J. Chem. Phys. 154, 074104 (2021); A. U. Bhatty and K. R. Brorsen, Mol. Phys. 119, e1936250 (2021)], where it was shown to provide accuracy on par with the most accurate vibrational structure methods with a low computational cost. Here, we eliminate the memory bottleneck of the second-order perturbation theory correction using the same (semi)stochastic approach developed previously for electronic structure theory. This allows us to treat, in an unbiased manner, much larger perturbative spaces, which are necessary for high accuracy in large systems. Stochastic errors are easily controlled to be less than 1 cm−1. We also report two other developments: (i) we propose a new heat-bath criterion and an associated exact implicit sorting algorithm for potential energy surfaces expressible as a sum of products of one-dimensional potentials; (ii) we formulate VHCI to use a vibrational self-consistent field (VSCF) reference, as opposed to the harmonic oscillator reference configuration used in previous reports. Our tests are done with quartic and sextic force fields, for which we find that with VSCF, the minor improvements to accuracy are outweighed by the higher computational cost associated the matrix element evaluations. We expect VSCF-based VHCI to be important for more general potential representations, for which the harmonic oscillator basis function integrals are no longer analytic.

Chemistry↗

Beyond Local Solvation Structure: Nanometric Aggregates in Battery Electrolytes and their Effect on Electrolyte Properties

Electrolytes are an essential component of all electrochemical storage and conversion devices, such as batteries. In the history of battery development, the complex nature of electrolytes has often been a bottleneck. Fundamental knowledge of electrolyte systems encompasses elucidation of structure-property relationships of the solution species. Recently, nanometric aggregates have been observed in several classes of electrolytes, including super-concentrated, redox-flow, multivalent, polymer, and ionic liquid-based electrolytes. Compared with the well-studied local solvation structures such as contact ion pairs and solvent-separated ions, these aggregates impose unique effects on the ion distribution and transport both within bulk electrolytes and at electrode/electrolyte interfaces. This Perspective highlights the discovery of the aggregates in various battery electrolytes and their impact on electrolyte properties. We also present an outlook for future studies of this emerging field of nanometric aggregates and the need for the development of new experimental and computational tools to study their properties.

Yu, Zhou↗

A Flexible Forwarding Scheme to Improve Latency-Bound Irregular P2P Communication in MPI

We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.

communication algorithm↗

Exploring structural transitions at grain boundaries in Nb using a generalized embedded atom interatomic potential

The advancement in experimental techniques, like the atom probe tomography and high resolution electron microscopy, is fueling interest in studying structural transformations of grain boundaries in metal and alloys to uncover correlations between mechanical properties and solute or impurity segregation to grain boundaries. Atomistic modeling is an important tool that can pinpoint the intricate dynamics of grain boundary phase transitions, but the lack of accurate interatomic potentials needed to simulate the complex dynamics of grain boundary structural transitions and identify different metastable phases has been the bottleneck. To this end, we use niobium as a model body centered cubic (BCC) metal and develop an interatomic potential to study grain boundary phase transitions. The potential for Nb is based on a generalization of the embedded atomic method potential and has sufficient flexibility to learn complex energy landscapes using a small set of training structures. We systematically test and validate the using data from ab initio density functional theory calculations and experiments. Using this potential, we calculate energies of multiple symmetric-tilt grain boundaries spanning a wide range of misorientation angles. Additionally, we explore different metastable structures of the Σ 27(552) [$1\overline{1}0$] grain boundary and use molecular dynamic simulations to study the coexistence of metastable phases and grain boundary transitions at finite temperature.

36 MATERIALS SCIENCE↗