Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Accelerator Design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix

Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adjacency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate 1.3x-110.5x improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.

Nair, Gopikrishnan R.↗

Community Toolkit for Designing and Implementing a Contractor Accelerator Program

This toolkit is designed to provide communities with an overview of the methodology and approach to designing, developing, and implementing contractor development programs. These programs support small businesses from historically under-represented communities so that they can better compete and become leaders in the clean energy marketplace. The toolkit was developed by Elevate and NREL for the Communities LEAP program.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Artificial Intelligence and Machine Learning for Bioenergy Research: Opportunities and Challenges

The integration of artificial intelligence and machine learning (AI/ML) with automated experimentation, genomics, biosystems design, and bioprocessing technologies is poised to revolutionize scientific investigation and, particularly, bioenergy research. To identify the opportunities and challenges in this emerging research area, the U.S. Department of Energy’s (DOE) Biological and Environmental Research program (BER) and Bioenergy Technologies Office (BETO) held a joint virtual workshop on AI/ML for Bioenergy Research (AMBER) on August 23–25, 2022. These interests have since been amplified in a September 2022 Executive Order, “Advancing Biotechnology and Biomanufacturing Innovation for a Sustainable, Safe, and Secure U.S. Bioeconomy,” to promote a whole-of government approach to biotechnology development (White House 2022). Approximately 50 scientists with various backgrounds and expertise from academia, industry, and DOE national laboratories met to discuss the opportunities and challenges of AI/ML for bioenergy research. Workshop participants were tasked with assessing the potential for AI/ML and laboratory automation to advance biological understanding and engineering in general. They particularly examined how integrating AI/ML tools with laboratory automation could accelerate biosystems design and optimize biomanufacturing. Discussions included the data and computational infrastructure needed to augment biosystems design applications and the expertise and workforce development efforts urgently required to shift integrated systems toward bioenergy research more broadly. Participants discussed many existing and future applications of AI/ML for biosystems design ranging from enzymes to plants and microbes, microbiomes, and bioprocess development. They also identified three key categories of scientific and technical opportunities and challenges: high-quality data, AI/ML algorithms, and laboratory automation. Several main takeaways emerged from the workshop: 1. Numerous AI/ML and automated experimentation applications exist for a variety of DOE mission needs in energy and the environment; 2. Exemplary research grand challenges for which AI/ML could provide solutions include: building microbes and microbial communities to specifications, developing closed-loop autonomous design and control for biosystems design, and advancing scale-up and automation; 3. Lack of sufficient high-quality, annotated data hinders the development of AI/ML applications; 4. New and improved AI/ML tools are needed, particularly those meeting the specific needs of the BER and BETO research communities; 5. Trade-offs in performance, cost, and reliability exist between deploying commercially available versus building custom-developed instrumentation and software for automated or autonomous experimentation; translation of manual to automated or autonomous methods is often a nontrivial endeavor; 6. Training a new generation of young scientists who can develop and apply AI/ML tools is needed to solve long-standing scientific challenges in bioenergy research. The integration of AI/ML tools and automated experimentation represents a new data-driven research paradigm complementary to the traditional hypothesis-driven research paradigm. This paradigm accelerates design and optimization of biological systems and processes for a variety of DOE mission needs in energy and the environment. The AMBER workshop broadly explored the potential of this new paradigm for bioenergy research, of particular interest to BER and BETO, and identified key challenges and opportunities that DOE can address in the coming years by leveraging its unique capabilities and resources.

59 BASIC BIOLOGICAL SCIENCES↗

Thermodynamic Modeling of the Al-Ce-Cu-Mg-Si System and Its Application to Aluminum-Cerium Alloy Design

Recently discovered AlCe alloys have shown promise in a number of applications, but the propensity of Ce to react with Al and other alloying elements can complicate the phase equilibria and design approach. To solve this, the CALPHAD method is used to explore an alloy within the quinary Al-Ce-Cu-Mg-Si system by developing a thermodynamic database with self-consistent parameters. The database includes a description of all 10 binary systems and 8 ternary systems consisting of: (i) 6 Al-containing ternaries (Al-Ce-Cu, Al-Ce-Mg, Al-Ce-Si, Al-Cu-Mg, Al-Cu-Si and Al-Mg-Si); and (ii) 2 additional ternaries that include Mg and Si (i.e., Ce-Mg-Si and Cu-Mg-Si). The thermodynamic description for the Al-Ce-Mg and Al-Mg-Si systems were reassessed to ensure consistency with the binary systems and the Ce-Mg-Si system is presented for the first time and compared to theoretical data from DFT (Density Functional Theory). In addition to the ternary interactions, the quaternary compound Al3Cu2Mg9Si7 and solid solution extending from the ternary Al2CuMg phase (Al,Si)2CuMg are incorporated. The CALPHAD method is employed and leveraged through the use of a Materials Design Simulator (MDS) to accelerate the design of novel aluminum-cerium-based alloys. The combination of a CALPHAD-based framework with experimental efforts and industrial insight permits the development of three new Al-Ce alloys: Al-3.5Ce-0.4Mg-7Si (Ce-modified A356), Al-5Ce-1Cu-0.5Mg-10Si and Al-19Ce-0.9Mg-1.1Si.

36 MATERIALS SCIENCE↗

Accelerating Traction Motor Optimization Design with AI Surrogate Models

The advancement of artificial intelligence systems enables the use of data-driven physics-based surrogate models to explore design spaces rapidly and deeply for engineering projects. This work presents a surrogate model workflow that accelerates electric traction motor design optimization by replacing finite element analysis (FEA) with an artificial neural network (ANN) and using this model in a genetic algorithm for design optimization. A baseline interior permanent-magnet motor is parameterized and sampled to generate FEA-labeled training data, after which a feed-forward ANN predicts key outputs (e.g., loss components and weight). The validated surrogate enables genetic-algorithm optimization and deep search over the design space without new FEA runs, producing Pareto-optimal trade-offs between weight and losses and set of optimized designs for rapid downselection of manufacturable motor designs.

Ribeiro, Pedro [ORNL] (ORCID:0009000921026641)↗

ML-AMD/exa-amd

ML-AMD is a Python workflow framework designed to accelerate the discovery and design of functional materials.

Moraru, Maxim [Los Alamos National Laboratory]↗

Multi-objective optimization with an integrated electromagnetics and beam dynamics workflow

In particle accelerators, RF cavities are used to accelerate charged particle beams to designed high energy for physical applications. In a typical accelerator design, the optimization of RF cavities and the optimization of beam dynamics are carried out in separate studies. For a more general and unrestricted accelerator design, a coupled optimization of the RF cavities and the beam parameters is required. For this coupled optimization problem, we have developed an integrated electromagnetics and beam dynamics workflow management system. Within this system, the geometries for a set of cavity components are first adjusted; the field modes are then computed with an electromagnetics program, and imported into a beam dynamics program for beam dynamics simulation. This workflow is encapsulated into a parallel multi-objective optimizer to achieve the integrated accelerator design optimization. A multi fidelity strategy is developed to improve the speed of the optimizer. Furthermore, this integrated global optimization capability is illustrated using a photoinjector design example and yields an improved design.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data-driven electrolyte design for lithium metal anodes

Improving Coulombic efficiency (CE) is key to the adoption of high energy density lithium metal batteries. Liquid electrolyte engineering has emerged as a promising strategy for improving the CE of lithium metal batteries, but its complexity renders the performance prediction and design of electrolytes challenging. Here, we develop machine learning (ML) models that assist and accelerate the design of high-performance electrolytes. Using the elemental composition of electrolytes as the features of our models, we apply linear regression, random forest, and bagging models to identify the critical features for predicting CE. Our models reveal that a reduction in the solvent oxygen content is critical for superior CE. We use the ML models to design electrolyte formulations with fluorine-free solvents that achieve a high CE of 99.70%. This work highlights the promise of data-driven approaches that can accelerate the design of high-performance electrolytes for lithium metal batteries.

25 ENERGY STORAGE↗

Preliminary Design of a LAMP DTL

A preliminary Drift Tube Linac (DTL) layout was designed to create an algorithm for developing the conceptual design of the proposed LANSCE Accelerator Modernization Project (LAMP) final section: the proton linear accelerator from 3 MeV to 100 MeV. Preceding reports describe the proposed layouts of the LEBT, the RFQ, and the MEBT subsections of the linac. Initial estimates of the needed RF power for the DTL are also presented in the report. Present report describes the initial layout of the DTL and the first longitudinal beam dynamics results of the simulations n simplified models. The follow-up reports will include the transverse focusing scheme and beam dynamics in details. Present report in the last section will only describe the planned focusing scheme, based on the existing LANCE DTL.

43 PARTICLE ACCELERATORS↗

Preliminary Design of a LAMP DTL (Update)

Present report is a continuation of the previously published report LA-UR-23-26600 from 19-Jun-2023 and incorporates the mentioned report as its part. All the additional information is contained in the part III of this report, the DTL Transverse Focusing Scheme. A preliminary Drift Tube Linac (DTL) layout was designed to create an algorithm for developing the conceptual design of the proposed LANSCE Accelerator Modernization Project (LAMP) final section: the proton linear accelerator from 3 MeV to 100 MeV. Preceding reports describe the proposed layouts of the LEBT, the RFQ, and the MEBT subsections of the linac. Initial estimates of the needed RF power for the DTL are also presented in the report. Present report describes the initial layout of the DTL and the first longitudinal beam dynamics (BD) results of the simulations in simplified models. The part III of this report includes the transverse focusing scheme and transverse BD in details.

43 PARTICLE ACCELERATORS↗

Advanced Modeling of Conventional Particle Accelerators

SciDAC-5 goals: Deliver particle accelerator and beam simulations tools that go beyond the current state of the art, up to the realization of virtual twins of particle accelerators, enabling design and modeling of particle accelerators at unprecedented speed, levels of accuracy, and realism; and apply these tools to key accelerator facilities relevant to DOE HEP (such as PIP-II/DUNE, FACET-II).

43 PARTICLE ACCELERATORS↗

Exploring the Use of Novel Spatial Accelerators in Scientific Applications

Driven by the need to find alternative accelerators which can viably replace GPUs in next-generation Supercomputing systems, this paper proposes a methodology to enable agile application/hardware co-design. The application-first methodology provides the ability to come up with design of accelerators while working with real-world workloads, available accelerators, and system software. The iterative design process targets a set of kernels in a workload for performance estimates that can prune the design space for later phases of detailed architectural evaluations. To this effect, in this paper, a novel data-parallel device model is introduced that simulates the latency of performance-sensitive operations in an accelerator including data transfers and kernel computation using multi-core CPUs. The use of off-the-shelf simulators, such as pre-RTL simulator Aladdin or multiple tools available for exploring the design of deep neural network accelerators (e.g., Timeloop) is demonstrated for evaluation of various accelerator designs using applications with realistic inputs. Examples of multiple device configurations that are instantiable in a system are explored to evaluate the performance benefit of deploying novel accelerators. The proposed device is integrated with a programming model and system software to potentially explore the impacts of high-level programming languages/compilers and low-level effects such as task scheduling on multiple accelerators. We analyze our methodology for a set of applications that represent high-performance computing (HPC) and graph analytics. The applications include a computational chemistry kernel realized using tensor contractions, triangle counting, GraphSAGE and Breadth-first Search. These applications include kernels such as dense matrix-dense matrix multiplication, sparse matrix-spare matrix multiplication, and sparse matrix-dense vector multiplication. Our results indicate potential performance benefits and insights for system design by including accelerators that realize these kernels along-side general purpose accelerators.

AI, codesign, Accelerated Computing, Modeling and ↗

Northstar Accelerator Based Mo99 Production Facility Design Support

Facility design support in FY20 centered on transfer of the target insertion and design details to MBraun and Springs, the companies under contract via Northstar to do the hot cell, target insertion and the local target shielding. A significant number of changes have been introduced by these vendors which have been reviewed and critiqued. In addition, the company contracted to produce the Mo100 disks announced that the tolerance on the thickness of 0.5 mm would be ±25%. An alternative design of the target holder was proposed. Bulk shielding and local target shielding design work has continued also and will be reported. Ongoing design and cooling issues continue to arise as the facility takes shape and moves toward completion. LANL is part of the team identifying, analyzing and designing solutions to these issues. Many of these are covered in this report.

43 PARTICLE ACCELERATORS↗

Accelerating Discovery of Solid‐State Thin‐Film Metal Dealloying for 3D Nanoarchitecture Materials Design through Laser Thermal Gradient Treatment

Thin‐film solid‐state metal dealloying (thin‐film SSMD) is a promising method for fabricating nanostructures with controlled morphology and efficiency, offering advantages over conventional bulk materials processing methods for integration into practical applications. Although machine learning (ML) has facilitated the design of dealloying systems, the selection of key thermal treatment parameters for nanostructure formation remains largely unknown and dependent on experimental trial and error. To overcome this challenge, a workflow enabling high‐throughput characterization of thermal treatment parameters is demonstrated using a laser‐based thermal treatment to create temperature gradients on single thin‐film samples of Nb‐Al/Sc and Nb‐Al/Cu. This continuous thermal space enables observation of dealloying transitions and the resulting nanostructures of interest. Through synchrotron X‐ray multimodal and high‐throughput characterization, critical transitions and nanostructures can be rapidly captured and subsequently verified using electron microscopy. The key temperatures driving chemical reactions and morphological evolutions are clearly identified. While the oxidation may influence nanostructure formation during thin‐film treatment, the dealloying process at the dealloying front involves interactions solely between the dealloying elements, highlighting the availability and viability of the selected systems. Further, this approach enables efficient exploration of the dealloying process and validation of ML predictions, thereby accelerating the discovery of thin‐film SSMD systems with targeted nanostructures.

36 MATERIALS SCIENCE↗

Language models for the prediction of SARS-CoV-2 inhibitors

The COVID-19 pandemic highlights the need for computational tools to automate and accelerate drug design for novel protein targets. We leverage deep learning language models to generate and score drug candidates based on predicted protein binding affinity. We pre-trained a deep learning language model (BERT) on ∼9.6 billion molecules and achieved peak performance of 603 petaflops in mixed precision. Our work reduces pre-training time from days to hours, compared to previous efforts with this architecture, while also increasing the dataset size by nearly an order of magnitude. For scoring, we fine-tuned the language model using an assembled set of thousands of protein targets with binding affinity data and searched for inhibitors of specific protein targets, SARS-CoV-2 Mpro and PLpro. We utilized a genetic algorithm approach for finding optimal candidates using the generation and scoring capabilities of the language model. Our generalizable models accelerate the identification of inhibitors for emerging therapeutic targets.

Blanchard, Andrew E.↗

Unveiling X-ray absorption signatures of boron nitride via first-principles simulation and machine learning

Boron nitride (BN) allotropes hold great promise in many advanced applications ranging from optical and photonic devices to energy storage and battery systems to tribological components. The diverse functionalities of this material stem from BN’s highly tunable structural and electronic properties, which are governed by the versatile boron–nitrogen bonding configurations. Exploring the structural landscape of BN can unveil novel structures possessing unique properties suited for specific applications, therefore accelerating the design of next-generation advanced functional materials. In this work, we leverage boron K-edge X-ray absorption spectroscopy (XAS) as an effective probe for local structural features and chemical environments. A total of 210 BN crystal structures are generated via analogies to the extensive array of carbon allotropes, and XAS is simulated for each unique local motif within the resulting collection of structures. A mapping between structural features and spectral signatures was established by synergizing first-principle simulations with data-driven based post-analysis approaches. Specifically, we developed a neural network model that can satisfactorily predict spectra line shapes from local structural descriptors. Toward automatic spectroscopic interpretation of any new BN structures, supervised machine learning models, trained on this structure–spectrum dataset, can accurately infer local coordination environments from simulated XAS, highlighting the strength of this unique approach of combining high-fidelity first-principles simulation and machine-learning to accelerate target design of novel BN materials via rational understanding of local structure-spectrum correlations.

36 MATERIALS SCIENCE↗