Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “library design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Co-design for Particle Applications at Exascale

Co-design across the Exascale Computing Project (ECP) has been critical for both enabling science applications and bringing disparate communities together. Developing and porting applications to the various high-performance computing (HPC) architectures on pre-exascale and exascale computers has been quite challenging due to the diversity of hardware features and software stacks. The Co-design Center for Particle Applications (CoPA) has developed and enhanced the Cabana and PROGRESS/BML libraries to facilitate the creation of new particle applications, make existing particle applications exascale capable, and allow teams to explore new capabilities. Particle methods from atomistic, mesoscale, continuum, through cosmological scales have been built with Cabana, along with new possibilities for application coupling. Similarly, the PROGRESS/BML library has enabled quantum particle applications with linear algebra solvers to use advanced hardware. Across these CoPA-developed libraries, the co-design abstraction layer combines performance portability with math library support to facilitate separation of concerns and directly support science runs.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Berkeley eXtensible Environment (BXE) v3

The Berkeley eXtensible Environment (BXE) provides a cloud environment for hardware designers and computer architects to design, build, and simulate their custom architectures on an on-premises FPGA cluster. Utilizing the Chipyard, MoSAIC, and FireSim frameworks, users are provided an environment where they can assemble SoC designs from an existing library of components or import their own source code. Once their designs are ready, they can utilize the FireSim framework provided by BXE to deploy and simulate their designs on the FPGA. Users aren't limited to a single FPGA; they can deploy multiple instances across multiple FPGAs, acting like a rack of servers, or partition their large design across multiple FPGAs, ganging multiple FPGAs into a single simulated system.

Fatollahi-Fard, Farzin↗

autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm Architectures

This paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPC-grade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM.

Wu, Du↗

RANGE: A robust adaptive nature-inspired global explorer of potential energy surfaces

With the growing demand for realistic representations of chemical structures and the advent of exascale computing, the intelligent sampling of potential energy surfaces and efficient identification of global minima have become more essential but also more feasible. Building on prior studies demonstrating the efficiency of the Artificial Bee Colony (ABC) swarm intelligence algorithm, we report a hybrid metaheuristic framework that integrates the adaptive exploration capabilities of ABC coupled with the exploitation strengths of genetic algorithms (GA) in a scalable, Python-based implementation. The resulting tool, RANGE (Robust Adaptive Nature-inspired Global Explorer), provides seamless interfaces to multiple potential energy evaluators, either directly or via widely used Python libraries, and is designed for high-performance computing environments. We describe the implementation details of RANGE and evaluate its performance, relative to ABC- or GA-alone based algorithms, on a variety of chemical systems, including molecular clusters and heterogeneous surfaces. In conclusion, our results demonstrate RANGE’s efficiency, robustness, and broad applicability in addressing challenging global optimization problems in computational chemistry and materials science.

Algorithms and data structure↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Superconducting Qubits Underground

Superconducting qubit devices are highly sensitive to cosmic ray induced quasiparticle poisoning—a challenge for quantum computing. At Fermilab, we utilize an underground testbed called QUIET that significantly decreases the rate of such events. We employ an open-source toolchain, including Qiskit Metal and the SQuADDs library, for device design and simulation prior to fabrication. We then test these chips in QUIET, which is capable of housing many chips, enabling tests of scalability. This presentation will cover our qubit design approach, fabrication capabilities, and first measurements from qubits deployed in QUIET, benchmarking the new facility.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-Throughput Microfluidic Electroporation (HTME): A Scalable, 384-Well Platform for Multiplexed Cell Engineering

Electroporation-mediated gene delivery is a cornerstone of synthetic biology, offering several advantages over other methods: higher efficiencies, broader applicability, and simpler sample preparation. Yet, electroporation protocols are often challenging to integrate into highly multiplexed workflows, owing to limitations in their scalability and tunability. These challenges ultimately increase the time and cost per transformation. As a result, rapidly screening genetic libraries, exploring combinatorial designs, or optimizing electroporation parameters requires extensive iterations, consuming large quantities of expensive custom-made DNA and cell lines or primary cells. To address these limitations, we have developed a High-Throughput Microfluidic Electroporation (HTME) platform that includes a 384-well electroporation plate (E-Plate) and control electronics capable of rapidly electroporating all wells in under a minute with individual control of each well. Fabricated using scalable and cost-effective printed-circuit-board (PCB) technology, the E-Plate significantly reduces consumable costs and reagent consumption by operating on nano to microliter volumes. Furthermore, individually addressable wells facilitate rapid exploration of large sets of experimental conditions to optimize electroporation for different cell types and plasmid concentrations/types. Use of the standard 384-well footprint makes the platform easily integrable into automated workflows, thereby enabling end-to-end automation. We demonstrate transformation of E. coli with pUC19 to validate the HTME's core functionality, achieving at least a single colony forming unit in more than 99% of wells and confirming the platform's ability to rapidly perform hundreds of electroporations with customizable conditions. This work highlights the HTME's potential to significantly accelerate synthetic biology Design-Build-Test-Learn (DBTL) cycles by mitigating the transformation/transfection bottleneck.

Gaillard, William R↗

Development of Masitinib Derivatives with Enhanced Mpro Ligand Efficiency and Reduced Cytotoxicity

Recently, a high-throughput screen of 1900 clinically used drugs identified masitinib, an orally bioavailable tyrosine kinase inhibitor, as a potential treatment for COVID-19. Masitinib acts as a broad-spectrum inhibitor for human coronaviruses, including SARS-CoV-2 and several of its variants. In this work, we rely on atomistic molecular dynamics simulations with advanced sampling methods to develop a deeper understanding of masitinib’s mechanism of M pro inhibition. To improve the inhibitory efficiency and to increase the ligand selectivity for the viral target, we determined the minimal portion of the molecule (fragment) that is responsible for most of the interactions that arise within the masitinib-Mpro complex. We found that masitinib forms highly stable and specific H-bond interactions with Mpro through its pyridine and aminothiazole rings. Importantly, the interaction with His 163 is a key anchoring point of the inhibitor, and its perturbation leads to ligand unbinding within nanoseconds. Based on these observations, a small library of rationally designed masitinib derivatives (M1–M5) was proposed. Our results show increased inhibitory efficiency and highly reduced cytotoxicity for the M3 and M4 derivatives compared to masitinib.

60 APPLIED LIFE SCIENCES↗

Optical tags comprising rare earth metal-organic frameworks

Optical tags provide a way to identify assets quickly and unambiguously, an application relevant to anti-counterfeiting and protection of valuable resources or information. The present invention is directed to a tag fluorophore that encodes multilayer complexity in a family of heterometallic rare-earth metal-organic frameworks (RE-MOFs) based on highly connected polynuclear clusters and carboxylic acid-based linkers. Both overt (visible) and covert (near infrared, NIR) properties with concomitant multi-emissive spectra and tunable luminescence lifetimes impart both intricacy and security. Tag authentication can be validated with a variety of orthogonal detection methodologies. The relationships between structure, composition, and optical properties of the family of RE-MOFs can be used to create a large library of rationally designed, highly complex, difficult to counterfeit optical tags.

Sava Gallis, Dorina F.↗

Graph Contractions for Calculating Correlation Functions in Lattice QCD

Computing correlation functions for many-particle systems in Lattice QCD is vital to extract nuclear physics observables like the energy spectrum of hadrons such as protons. However, this type of calculation has long been considered to be very challenging and computing-resource intensive because of the complex nature of a hadron composed of quarks with many degrees of freedom. In particular, a correlation function can be calculated through a sum of all possible pairs of quark contractions, each of which is a batched tensor contraction, dictated by Wick's theorem. Because the number of terms of this sum can be very large for any hadronic system of interest, fast evaluation of the sum faces several challenges: an extremely large number of contractions, a huge memory footprint at runtime, and the speed of tensor contractions. In this paper, we present a Lattice QCD analysis software suite, Redstar, which addresses these challenges by utilizing novel algorithmic and software engineering methods targeting modern computing platforms such as many-core CPUs and GPUs. In particular, Redstar represents every term in the sum of a correlation function by a graph, applies efficient graph algorithms to reduce the number of contractions to lower the cost of computations, and minimizes the total memory footprint. Moreover, Redstar carries out the contractions on either CPUs or GPUs utilizing an internal and highly efficient Hadron contraction library. Specifically, we illustrate some important algorithmic optimizations of Redstar, show various key design features of Hadron library, and present the speedup values due to the optimizations along with performance figures for calculating six correlations functions on four computing platforms.

Chen, Jie↗

Graph Contractions for Calculating Correlation Functions in Lattice QCD

Computing correlation functions for many-particle systems in Lattice QCD is vital to extract nuclear physics observables like the energy spectrum of hadrons such as protons. However, this type of calculation has long been considered to be very challenging and computing-resource intensive because of the complex nature of a hadron composed of quarks with many degrees of freedom. In particular, a correlation function can be calculated through a sum of all possible pairs of quark contractions, each of which is a batched tensor contraction, dictated by Wick's theorem. Because the number of terms of this sum can be very large for any hadronic system of interest, fast evaluation of the sum faces several challenges: an extremely large number of contractions, a huge memory footprint at runtime, and the speed of tensor contractions. In this paper, we present a Lattice QCD analysis software suite, Redstar, which addresses these challenges by utilizing novel algorithmic and software engineering methods targeting modern computing platforms such as many-core CPUs and GPUs. In particular, Redstar represents every term in the sum of a correlation function by a graph, applies efficient graph algorithms to reduce the number of contractions to lower the cost of computations, and minimizes the total memory footprint. Moreover, Redstar carries out the contractions on either CPUs or GPUs utilizing an internal and highly efficient Hadron contraction library. Specifically, we illustrate some important algorithmic optimizations of Redstar, show various key design features of Hadron library, and present the speedup values due to the optimizations along with performance figures for calculating six correlations functions on four computing platforms.

Chen, Jie↗

A single-atom library for guided monometallic and concentration-complex multimetallic designs

Atomically dispersed single-atom catalysts have the potential to bridge heterogeneous and homogeneous catalysis. Dozens of single-atom catalysts have been developed, and they exhibit notable catalytic activity and selectivity that are not achievable on metal surfaces. Although promising, there is limited knowledge about the boundaries for the monometallic single-atom phase space, not to mention multimetallic phase spaces. Here, single-atom catalysts based on 37 monometallic elements are synthesized using a dissolution-and-carbonization method, characterized and analyzed to build the largest reported library of single-atom catalysts. In conjunction with in situ studies, we uncover unified principles on the oxidation state, coordination number, bond length, coordination element and metal loading of single atoms to guide the design of single-atom catalysts with atomically dispersed atoms anchored on N-doped carbon. We utilize the library to open up complex multimetallic phase spaces for single-atom catalysts and demonstrate that there is no fundamental limit on using single-atom anchor sites as structural units to assemble concentration-complex single-atom catalyst materials with up to 12 different elements. Furthermore, our work offers a single-atom library spanning from monometallic to concentration-complex multimetallic materials for the rational design of single-atom catalysts.

36 MATERIALS SCIENCE↗

ReaLigands: A Ligand Library Cultivated from Experiment and Intended for Molecular Computational Catalyst Design

Computational catalyst design requires identification of a metal and ligand that together result in the desired reaction reactivity and/or selectivity. A major impediment to translating computational designs to experiments is evaluating ligands that are likely to be synthesized. Here we provide a solution to this impediment with our ReaLigands library that contains >30,000 monodentate, bidentate (didentate), tridentate, and larger ligands cultivated by dismantling experimentally reported crystal structures. Individual ligands from mononuclear crystal structures were identified using a modified depth-first search algorithm and charge was assigned using a machine learning model based on quantum-chemical calculated features. In the library ligands are sorted based on direct ligand-to-metal atomic connections and on denticity. Representative principal component analysis (PCA) and uniform manifold approximation and projection (UMAP) analyses were used to analyze several tridentate ligand categories, which revealed both the diversity of ligands and connections between ligand categories. Furthermore, we also demonstrated the utility of this library by implementing it with our building and optimization tools, which resulted in the very rapid generation of barriers for 750 bidentate ligands for Rh-hydride ethylene migratory insertion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

De novo designed protein inhibitors of amyloid aggregation and seeding

Neurodegenerative diseases are characterized by the pathologic accumulation of aggregated proteins. Known as amyloid, these fibrillar aggregates include proteins such as tau and amyloid-β (Aβ) in Alzheimer’s disease (AD) and alpha-synuclein (αSyn) in Parkinson’s disease (PD). The development and spread of amyloid fibrils within the brain correlates with disease onset and progression, and inhibiting amyloid formation is a possible route toward therapeutic development. Recent advances have enabled the determination of amyloid fibril structures to atomic-level resolution, improving the possibility of structure-based inhibitor design. In this work, we use these amyloid structures to design inhibitors that bind to the ends of fibrils, “capping” them so as to prevent further growth. Using de novo protein design, we develop a library of miniprotein inhibitors of 35 to 48 residues that target the amyloid structures of tau, Aβ, and αSyn. Biophysical characterization of top in silico designed inhibitors shows they form stable folds, have no sequence similarity to naturally occurring proteins, and specifically prevent the aggregation of their targeted amyloid-prone proteins in vitro. The inhibitors also prevent the seeded aggregation and toxicity of fibrils in cells. In vivo evaluation reveals their ability to reduce aggregation and rescue motor deficits in Caenorhabditis elegans models of PD and AD.

59 BASIC BIOLOGICAL SCIENCES↗

REYNO: A Reactive Hydrodynamics Modeling Suite

REYNO is a Python library that is primarily designed as a platform for the agile development and testing of novel reactive flow models. Equations of state and multi-step rate laws can be implemented using the provided class hierarchy with different thermodynamic closure conditions. The library includes a one-dimensional Lagrangian hydrodynamic solver to simulate multi-material reactive or inert problems in planar, cylindrical, or spherical coordinates. Custom time-dependent boundary conditions are also available for problems such as ramp compression. This document describes the different numerical methods and algorithms that are included in REYNO.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Thermophilic Chassis-Enabled High-Throughput Selection of a Thermostable Fluorogenic Reporter

Thermostable proteins show increased shelf life and performance at elevated temperatures and under harsh conditions, resulting in lower costs for various industrial and biotechnological applications. However, due to a limited understanding of the relationship between stability and function, protein stabilization remains primarily a trial-and-error approach. Therefore, building a combinatorial library of mutations predicted to improve stability, followed by experimental testing, represents a markedly improved methodology. However, the lack of high-throughput approaches to screen even a moderately sized library presents a major bottleneck in the field. Here, in this study, we use a thermophile, Parageobacillus thermoglucosidasius (Ptherm) to rapidly screen combinatorial libraries consisting of rationally designed thermostabilizing mutations (∼10 3 –10 4 ) of a mesophilic fluorescent reporter, Y-FAST. On a Petri dish, microbial growth at an elevated temperature and exposure to fluorogen yielded several colonies of Ptherm that showed distinct fluorescence at 55 and 68 °C in our two sequentially generated libraries using Rosetta and ProteinMPNN, respectively. The Y-FAST variants isolated from fluorescent colonies were brighter than Y-FAST and showed higher resistance to thermal and chemical denaturation. AlphaFold-predicted structures and MD simulations revealed stability-enhancing salt bridges and hydrogen bond networks in the isolated FAST variants. The moderately thermostable FAST (tsFAST) and hyperstable FAST (hsFAST) were then demonstrated as translation reporters for protein expression and folding at elevated temperatures, such as 55 and 68 °C. Our approach of combinatorial library generation and high-throughput screening in a thermophilic chassis could, in principle, be extended to other proteins fused to these translation reporters. Furthermore, the hsFAST protein is small─half the size of the green fluorescent protein─and does not require oxygen for maturation, making it ideal for engineering extremophilic anaerobes for biosensing and bioconversion.

59 BASIC BIOLOGICAL SCIENCES↗