Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “library design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

A Robust and Scalable Software Library for Parallel Adaptive Refinement on Unstructured Meshes

The design and implementation of Pyramid, a software library for performing parallel adaptive mesh refinement (PAMR) on unstructured meshes, is described. This software library can be easily used in a variety of unstructured parallel computational applications, including parallel finite element, parallel finite volume, and parallel visualization applications using triangular or tetrahedral meshes. The library contains a suite of well-designed and efficiently implemented modules that perform operations in a typical PAMR process. Among these are mesh quality control during successive parallel adaptive refinement (typically guided by a local-error estimator), parallel load-balancing, and parallel mesh partitioning using the ParMeTiS partitioner. The Pyramid library is implemented in Fortran 90 with an interface to the Message-Passing Interface (MPI) library, supporting code efficiency, modularity, and portability. An EM waveguide filter application, adaptively refined using the Pyramid library, is illustrated.

Lou, John Z.↗

Berkeley eXtensible Environment (BXE) v3

The Berkeley eXtensible Environment (BXE) provides a cloud environment for hardware designers and computer architects to design, build, and simulate their custom architectures on an on-premises FPGA cluster. Utilizing the Chipyard, MoSAIC, and FireSim frameworks, users are provided an environment where they can assemble SoC designs from an existing library of components or import their own source code. Once their designs are ready, they can utilize the FireSim framework provided by BXE to deploy and simulate their designs on the FPGA. Users aren't limited to a single FPGA; they can deploy multiple instances across multiple FPGAs, acting like a rack of servers, or partition their large design across multiple FPGAs, ganging multiple FPGAs into a single simulated system.

Fatollahi-Fard, Farzin↗

autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm Architectures

This paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPC-grade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM.

Wu, Du↗

RANGE: A robust adaptive nature-inspired global explorer of potential energy surfaces

With the growing demand for realistic representations of chemical structures and the advent of exascale computing, the intelligent sampling of potential energy surfaces and efficient identification of global minima have become more essential but also more feasible. Building on prior studies demonstrating the efficiency of the Artificial Bee Colony (ABC) swarm intelligence algorithm, we report a hybrid metaheuristic framework that integrates the adaptive exploration capabilities of ABC coupled with the exploitation strengths of genetic algorithms (GA) in a scalable, Python-based implementation. The resulting tool, RANGE (Robust Adaptive Nature-inspired Global Explorer), provides seamless interfaces to multiple potential energy evaluators, either directly or via widely used Python libraries, and is designed for high-performance computing environments. We describe the implementation details of RANGE and evaluate its performance, relative to ABC- or GA-alone based algorithms, on a variety of chemical systems, including molecular clusters and heterogeneous surfaces. In conclusion, our results demonstrate RANGE’s efficiency, robustness, and broad applicability in addressing challenging global optimization problems in computational chemistry and materials science.

Algorithms and data structure↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Superconducting Qubits Underground

Superconducting qubit devices are highly sensitive to cosmic ray induced quasiparticle poisoning—a challenge for quantum computing. At Fermilab, we utilize an underground testbed called QUIET that significantly decreases the rate of such events. We employ an open-source toolchain, including Qiskit Metal and the SQuADDs library, for device design and simulation prior to fabrication. We then test these chips in QUIET, which is capable of housing many chips, enabling tests of scalability. This presentation will cover our qubit design approach, fabrication capabilities, and first measurements from qubits deployed in QUIET, benchmarking the new facility.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-Throughput Microfluidic Electroporation (HTME): A Scalable, 384-Well Platform for Multiplexed Cell Engineering

Electroporation-mediated gene delivery is a cornerstone of synthetic biology, offering several advantages over other methods: higher efficiencies, broader applicability, and simpler sample preparation. Yet, electroporation protocols are often challenging to integrate into highly multiplexed workflows, owing to limitations in their scalability and tunability. These challenges ultimately increase the time and cost per transformation. As a result, rapidly screening genetic libraries, exploring combinatorial designs, or optimizing electroporation parameters requires extensive iterations, consuming large quantities of expensive custom-made DNA and cell lines or primary cells. To address these limitations, we have developed a High-Throughput Microfluidic Electroporation (HTME) platform that includes a 384-well electroporation plate (E-Plate) and control electronics capable of rapidly electroporating all wells in under a minute with individual control of each well. Fabricated using scalable and cost-effective printed-circuit-board (PCB) technology, the E-Plate significantly reduces consumable costs and reagent consumption by operating on nano to microliter volumes. Furthermore, individually addressable wells facilitate rapid exploration of large sets of experimental conditions to optimize electroporation for different cell types and plasmid concentrations/types. Use of the standard 384-well footprint makes the platform easily integrable into automated workflows, thereby enabling end-to-end automation. We demonstrate transformation of E. coli with pUC19 to validate the HTME's core functionality, achieving at least a single colony forming unit in more than 99% of wells and confirming the platform's ability to rapidly perform hundreds of electroporations with customizable conditions. This work highlights the HTME's potential to significantly accelerate synthetic biology Design-Build-Test-Learn (DBTL) cycles by mitigating the transformation/transfection bottleneck.

Gaillard, William R↗

Development of Masitinib Derivatives with Enhanced Mpro Ligand Efficiency and Reduced Cytotoxicity

Recently, a high-throughput screen of 1900 clinically used drugs identified masitinib, an orally bioavailable tyrosine kinase inhibitor, as a potential treatment for COVID-19. Masitinib acts as a broad-spectrum inhibitor for human coronaviruses, including SARS-CoV-2 and several of its variants. In this work, we rely on atomistic molecular dynamics simulations with advanced sampling methods to develop a deeper understanding of masitinib’s mechanism of M pro inhibition. To improve the inhibitory efficiency and to increase the ligand selectivity for the viral target, we determined the minimal portion of the molecule (fragment) that is responsible for most of the interactions that arise within the masitinib-Mpro complex. We found that masitinib forms highly stable and specific H-bond interactions with Mpro through its pyridine and aminothiazole rings. Importantly, the interaction with His 163 is a key anchoring point of the inhibitor, and its perturbation leads to ligand unbinding within nanoseconds. Based on these observations, a small library of rationally designed masitinib derivatives (M1–M5) was proposed. Our results show increased inhibitory efficiency and highly reduced cytotoxicity for the M3 and M4 derivatives compared to masitinib.

60 APPLIED LIFE SCIENCES↗

Arabidopsis chloroplast chaperonin 10 is a calmodulin-binding protein

Calcium regulates diverse cellular activities in plants through the action of calmodulin (CaM). By using (35)S-labeled CaM to screen an Arabidopsis seedling cDNA expression library, a cDNA designated as AtCh-CPN10 (Arabidopsis thaliana chloroplast chaperonin 10) was cloned. Chloroplast CPN10, a nuclear-encoded protein, is a functional homolog of E. coli GroES. It is believed that CPN60 and CPN10 are involved in the assembly of Rubisco, a key enzyme involved in the photosynthetic pathway. Northern analysis revealed that AtCh-CPN10 is highly expressed in green tissues. The recombinant AtCh-CPN10 binds to CaM in a calcium-dependent manner. Deletion mutants revealed that there is only one CaM-binding site in the last 31 amino acids of the AtCh-CPN10 at the C-terminal end. The CaM-binding region in AtCh-CPN10 has higher homology to other chloroplast CPN10s in comparison to GroES and mitochondrial CPN10s, suggesting that CaM may only bind to chloroplast CPN10s. Furthermore, the results also suggest that the calcium/CaM messenger system is involved in regulating Rubisco assembly in the chloroplast, thereby influencing photosynthesis. Copyright 2000 Academic Press.

NASA Discipline Plant Biology↗

Optical tags comprising rare earth metal-organic frameworks

Optical tags provide a way to identify assets quickly and unambiguously, an application relevant to anti-counterfeiting and protection of valuable resources or information. The present invention is directed to a tag fluorophore that encodes multilayer complexity in a family of heterometallic rare-earth metal-organic frameworks (RE-MOFs) based on highly connected polynuclear clusters and carboxylic acid-based linkers. Both overt (visible) and covert (near infrared, NIR) properties with concomitant multi-emissive spectra and tunable luminescence lifetimes impart both intricacy and security. Tag authentication can be validated with a variety of orthogonal detection methodologies. The relationships between structure, composition, and optical properties of the family of RE-MOFs can be used to create a large library of rationally designed, highly complex, difficult to counterfeit optical tags.

Sava Gallis, Dorina F.↗

Graph Contractions for Calculating Correlation Functions in Lattice QCD

Computing correlation functions for many-particle systems in Lattice QCD is vital to extract nuclear physics observables like the energy spectrum of hadrons such as protons. However, this type of calculation has long been considered to be very challenging and computing-resource intensive because of the complex nature of a hadron composed of quarks with many degrees of freedom. In particular, a correlation function can be calculated through a sum of all possible pairs of quark contractions, each of which is a batched tensor contraction, dictated by Wick's theorem. Because the number of terms of this sum can be very large for any hadronic system of interest, fast evaluation of the sum faces several challenges: an extremely large number of contractions, a huge memory footprint at runtime, and the speed of tensor contractions. In this paper, we present a Lattice QCD analysis software suite, Redstar, which addresses these challenges by utilizing novel algorithmic and software engineering methods targeting modern computing platforms such as many-core CPUs and GPUs. In particular, Redstar represents every term in the sum of a correlation function by a graph, applies efficient graph algorithms to reduce the number of contractions to lower the cost of computations, and minimizes the total memory footprint. Moreover, Redstar carries out the contractions on either CPUs or GPUs utilizing an internal and highly efficient Hadron contraction library. Specifically, we illustrate some important algorithmic optimizations of Redstar, show various key design features of Hadron library, and present the speedup values due to the optimizations along with performance figures for calculating six correlations functions on four computing platforms.

Chen, Jie↗

Graph Contractions for Calculating Correlation Functions in Lattice QCD

Computing correlation functions for many-particle systems in Lattice QCD is vital to extract nuclear physics observables like the energy spectrum of hadrons such as protons. However, this type of calculation has long been considered to be very challenging and computing-resource intensive because of the complex nature of a hadron composed of quarks with many degrees of freedom. In particular, a correlation function can be calculated through a sum of all possible pairs of quark contractions, each of which is a batched tensor contraction, dictated by Wick's theorem. Because the number of terms of this sum can be very large for any hadronic system of interest, fast evaluation of the sum faces several challenges: an extremely large number of contractions, a huge memory footprint at runtime, and the speed of tensor contractions. In this paper, we present a Lattice QCD analysis software suite, Redstar, which addresses these challenges by utilizing novel algorithmic and software engineering methods targeting modern computing platforms such as many-core CPUs and GPUs. In particular, Redstar represents every term in the sum of a correlation function by a graph, applies efficient graph algorithms to reduce the number of contractions to lower the cost of computations, and minimizes the total memory footprint. Moreover, Redstar carries out the contractions on either CPUs or GPUs utilizing an internal and highly efficient Hadron contraction library. Specifically, we illustrate some important algorithmic optimizations of Redstar, show various key design features of Hadron library, and present the speedup values due to the optimizations along with performance figures for calculating six correlations functions on four computing platforms.

Chen, Jie↗

A single-atom library for guided monometallic and concentration-complex multimetallic designs

Atomically dispersed single-atom catalysts have the potential to bridge heterogeneous and homogeneous catalysis. Dozens of single-atom catalysts have been developed, and they exhibit notable catalytic activity and selectivity that are not achievable on metal surfaces. Although promising, there is limited knowledge about the boundaries for the monometallic single-atom phase space, not to mention multimetallic phase spaces. Here, single-atom catalysts based on 37 monometallic elements are synthesized using a dissolution-and-carbonization method, characterized and analyzed to build the largest reported library of single-atom catalysts. In conjunction with in situ studies, we uncover unified principles on the oxidation state, coordination number, bond length, coordination element and metal loading of single atoms to guide the design of single-atom catalysts with atomically dispersed atoms anchored on N-doped carbon. We utilize the library to open up complex multimetallic phase spaces for single-atom catalysts and demonstrate that there is no fundamental limit on using single-atom anchor sites as structural units to assemble concentration-complex single-atom catalyst materials with up to 12 different elements. Furthermore, our work offers a single-atom library spanning from monometallic to concentration-complex multimetallic materials for the rational design of single-atom catalysts.

36 MATERIALS SCIENCE↗

An integrated runtime and compile-time approach for parallelizing structured and block structured applications

Scientific and engineering applications often involve structured meshes. These meshes may be nested (for multigrid codes) and/or irregularly coupled (called multiblock or irregularly coupled regular mesh problems). A combined runtime and compile-time approach for parallelizing these applications on distributed memory parallel machines in an efficient and machine-independent fashion was described. A runtime library which can be used to port these applications on distributed memory machines was designed and implemented. The library is currently implemented on several different systems. To further ease the task of application programmers, methods were developed for integrating this runtime library with compilers for HPK-like parallel programming languages. How this runtime library was integrated with the Fortran 90D compiler being developed at Syracuse University is discussed. Experimental results to demonstrate the efficacy of our approach are presented. A multiblock Navier-Stokes solver template and a multigrid code were experimented with. Our experimental results show that our primitives have low runtime communication overheads. Further, the compiler parallelized codes perform within 20 percent of the code parallelized by manually inserting calls to the runtime library.

Agrawal, Gagan↗

ReaLigands: A Ligand Library Cultivated from Experiment and Intended for Molecular Computational Catalyst Design

Computational catalyst design requires identification of a metal and ligand that together result in the desired reaction reactivity and/or selectivity. A major impediment to translating computational designs to experiments is evaluating ligands that are likely to be synthesized. Here we provide a solution to this impediment with our ReaLigands library that contains >30,000 monodentate, bidentate (didentate), tridentate, and larger ligands cultivated by dismantling experimentally reported crystal structures. Individual ligands from mononuclear crystal structures were identified using a modified depth-first search algorithm and charge was assigned using a machine learning model based on quantum-chemical calculated features. In the library ligands are sorted based on direct ligand-to-metal atomic connections and on denticity. Representative principal component analysis (PCA) and uniform manifold approximation and projection (UMAP) analyses were used to analyze several tridentate ligand categories, which revealed both the diversity of ligands and connections between ligand categories. Furthermore, we also demonstrated the utility of this library by implementing it with our building and optimization tools, which resulted in the very rapid generation of barriers for 750 bidentate ligands for Rh-hydride ethylene migratory insertion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

De novo designed protein inhibitors of amyloid aggregation and seeding

Neurodegenerative diseases are characterized by the pathologic accumulation of aggregated proteins. Known as amyloid, these fibrillar aggregates include proteins such as tau and amyloid-β (Aβ) in Alzheimer’s disease (AD) and alpha-synuclein (αSyn) in Parkinson’s disease (PD). The development and spread of amyloid fibrils within the brain correlates with disease onset and progression, and inhibiting amyloid formation is a possible route toward therapeutic development. Recent advances have enabled the determination of amyloid fibril structures to atomic-level resolution, improving the possibility of structure-based inhibitor design. In this work, we use these amyloid structures to design inhibitors that bind to the ends of fibrils, “capping” them so as to prevent further growth. Using de novo protein design, we develop a library of miniprotein inhibitors of 35 to 48 residues that target the amyloid structures of tau, Aβ, and αSyn. Biophysical characterization of top in silico designed inhibitors shows they form stable folds, have no sequence similarity to naturally occurring proteins, and specifically prevent the aggregation of their targeted amyloid-prone proteins in vitro. The inhibitors also prevent the seeded aggregation and toxicity of fibrils in cells. In vivo evaluation reveals their ability to reduce aggregation and rescue motor deficits in Caenorhabditis elegans models of PD and AD.

59 BASIC BIOLOGICAL SCIENCES↗