Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “accelerator modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Bayesian model updating with finite element vs surrogate models: Application to a miter gate structural system

Bayesian finite element (FE) model updating using direct model evaluations of large-scale high-fidelity FE models is extremely computationally expensive. Surrogate models can be used as fast emulators of FE models to accelerate the model calibration process. The physics/mechanics-based FE models are still the underpinning behind the surrogate models. Here, this paper evaluates the loss in accuracy and the gain in computational time while performing Bayesian model updating by using surrogate model evaluations compared to using direct FE model evaluations. This evaluation is crucial before entirely relying on surrogate models in model updating for structural health monitoring (SHM) and damage prognosis (DP) purposes. This paper also demonstrates Bayesian updating and surrogate model construction of large-scale high-fidelity FE models of infrastructure systems. In this regard, the miter gate structural system is considered as the testbed structure. Three predominant damage modes (loss of contact between gate and wall, loss of thickness due to corrosion, and loss of tension in the diagonal rods) are considered for model updating purposes. Bayesian model updating is performed using direct FE evaluations by leveraging parallel computing. Two types of surrogates, namely polynomial chaos expansion (PCE) and Gaussian process regression (GPR), are developed for the miter gate. Model updating is performed again using the trained surrogate models, and the updating results are compared with their counterparts obtained using the direct FE evaluation results. The posterior distribution of the FE model parameters obtained using the trained surrogates are sufficiently accurate with respect to the posterior obtained utilizing the direct FE evaluations. In addition, an approximate 4-fold decrease in the computational time was observed when using surrogate model evaluations instead of direct FE evaluations for model updating.

42 ENGINEERING↗

Neural-network accelerated coupled core-pedestal simulations with self-consistent transport of impurities and compatible with ITER IMAS

pedestal structure, current profile, and plasma equilibrium physics has been developed and tested against a DIII-D discharge. Here, key features of the achieved core-pedestal coupled workflow are its ability to account for the transport of impurities in the plasma self-consistently, as well as its use of machine learning accelerated models for the pedestal structure and for the turbulent transport physics. Notably, the coupled workflow is implemented within the OMFIT framework, and makes use of the ITER integrated modeling and analysis suite (IMAS) data structure for exchanging data among the physics codes that are involved in the simulations. Such technical advance has been facilitated by the development of a new numerical library named OMAS.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Framework for Neural Network Inference on FPGA-Centric SmartNICs

FPGA-based SmartNICs offer great potential to significantly improve the performance of high-performance computing and warehouse data processing by tightly coupling support for reconfigurable data-intensive computation with cross-node communication, thereby mitigating the von Neumann bottleneck. Existing work, however, has been generally been limited in that it assumes an accelerator model where kernels are offloaded to SmartNICs, but most control tasks are left to the CPUs. This leads to frequent waiting, inferior performance, and scaling challenges. In this work, we propose a new distributive data-centric computing framework, named FCsN, for reconfigurable SmartNIC-based systems. Through a lightweight task circulation execution model and its implementation architecture, FCsN allows the complete detaching of kernel execution, control logic, system scheduling, and network communication to the SmarNICs. This boosts performance by: (i) avoiding the control dependency with CPUs and (ii) supporting streaming kernel execution and network communication at line rate and in a very fine-grained manner. We demonstrate the efficiency and flexibility of FCsN using various types of neural network applications including graph neural networks; as these last are both irregular and data intensive they offer an especially robust demonstration. Evaluations using commonly-used neural network models and graph datasets show that a system with the support of FCsN can achieve, on average, 144 speedups over the MPI-based standard CPU baselines.

Guo, Anqi↗

Multi-Probe 24A Pre-Shot report

The Multi-Probe Radiography campaign will field a full day of experiments at the Omega EP laser facility. The day consists of shots alternating between Omega EP’s two short-pulse laser beams. The Backlighter beam will generate proton and deuteron beams from 500-800nm CH/CD film targets, and the Sidelighter beam will accelerate electrons to generate x-rays from 0.5mm x 25-50µm Ø Ta wire targets attached to Compound Parabolic Concentrator (CPC) cones. The ion beams shots will optimize CD foil thickness (maximizing ion energy/yield) for the transparency regime for use as the pitcher in a pitcher catcher neutron-generation platform. Several shots will include a LiF catcher within the NTA and this feeds into future neutron radiography work. X-ray shots will be used to characterize the platform established in FY2020B via measurement of electron spectra with and without the Ta wire, and radiography will be conducted at varying laser energies. This will validate PIC electron acceleration modeling while also demonstrating potential for using the x-ray platform for imaging of thicker objects.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Seeing is Believing: Autonomous Microscopy and the Data Revolution in Materials Science [Slides]

Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.

36 MATERIALS SCIENCE↗

γ-Ray Emission from Classical Nova V392 Per: Measurements from Fermi and HAWC

This paper reports on the γ -ray properties of the 2018 Galactic nova V392 Per, spanning photon energies ~0.1 GeV–100 TeV by combining observations from the Fermi Gamma-ray Space Telescope and the HAWC Observatory. As one of the most rapidly evolving γ -ray signals yet observed for a nova, GeV γ -rays with a power-law spectrum with an index Γ = 2.0 ± 0.1 were detected over 8 days following V392 Per’s optical maximum. HAWC observations constrain the TeV γ -ray signal during this time and also before and after. We observe no statistically significant evidence of TeV γ -ray emission from V392 Per, but present flux limits. Tests disfavor the extension of the Fermi Large Area Telescope spectrum to energies above 5 TeV by 2 standard deviations (95%) or more. We fit V392 Per’s GeV γ -rays with hadronic acceleration models, incorporating optical observations, and compare the calculations with HAWC limits.

79 ASTRONOMY AND ASTROPHYSICS↗

Interconnect: Cooperative Research and Development Final Report, CRADA Number CRD-13-00507 (Project 4)

This CRADA modification involves analyses on a variety of CdTe-PV related materials, test structures, solar cells, and modules produced at FSLR and/or NLR and adds analysis related to module degradation and reliability. Materials will be provided by FSLR, NLR, and/or by interleaving FSLR and NLR layers and processes. Module reliability activities will include technical risk assessment, materials characterization, module and test structure characterization, modeling, accelerated test development, and outdoor testing.

14 SOLAR ENERGY↗

The Rise of Intelligent Materials Science: Unleashing the Power of Machine Intelligence in Characterization

Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.

autonomous↗

Beyond Human Vision: Exploring Materials with Machine Intelligence

Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.

artificial intelligence↗

Machine learning accelerated discrete element modeling of granular flows

Granular flows are widely encountered in many industrial processes and natural phenomena. Discrete Element Modeling (DEM) is a useful tool for understanding and troubleshooting devices, which handle granular materials. However, its applicability is significantly limited by the huge computational cost associated with detecting and computing collisions. In this research, the computation speed of DEM was accelerated by orders of magnitude using a convolutional neural network to replace the direct calculation of particle-particle and particle-boundary collisions. The MFiX software was used to generate the training and testing dataset. Additionally, a GPU accelerated TensorFlow model was used to train the neural network and test the results. The model fluctuations caused by different training steps were reduced with a multi-scale loss function. The accuracy was improved with more frames within one training step. The modeling of a rotating drum and a hopper demonstrated the accuracy and efficiency of this machine learning accelerated DEM in the simulation of granular flows.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predictive Models and Novel Accelerated Tests for the Reliability of Cell Metallization in Photovoltaic Modules (Final Report)

Studies of metallization corrosion m photovoltaics have mainly been limited to comparisons of modules placed in accelerated chambers to fielded modules [1]. Damp Heat accelerated tests and phenomenological equations [2] are used to assess metallization corrosion without understanding the effect of UV light and temperature and humidity cycles on encapsulant adhesion degradation. The roles of encapsulant in-and out-diffusions of moisture and encapsulant impurities are important. Furthermore, few photovoltaic metallization corrosion studies included the role of bias and leakage currents, which are crucial in the electrochemical reaction of metallization. Leakage currents can highly accelerate the corrosion mechanism and are important to include in the studies of corrosion. In our research plan, we will address the following gaps in the PV community's understanding of metallization corrosion: (1) metallization corrosion with bias, humidity, and impurities in the encapsulant or metallization; (2) humidity diffusion through fresh and degraded encapsulants; and (3) comparison of model predictions with outdoor field modules and SunPower's extensive data for its back-contact and front-contact fleets [2] along with NREL's store of >20 year old modules. Our goal is to build models and accelerated tests to predict long term degradation of metallization corrosion of photovoltaic modules in the field. Our studies will include metals used in c-Si solar cells (Cu, Ag, and Al) and commonly used encapsulants (EV A (ethylene vinyl acetate), TPO (thermoplastic olefin), and silicone).

14 SOLAR ENERGY↗

Accelerating Adoption of Energy-Efficient Technology with the Thermalize Model

The accelerating impacts of climate change destabilize food systems and ecosystems alongside the communities that rely on them. Thus, climate resilience is a health equity issue at both a physical (relating to biological impacts of pollution and environmental degradation) and a cultural level (loss of subsistence lifestyles, languages, etc). With roughly 20% of US energy-related greenhouse gas emissions stemming from heating, cooling, and powering households, accessible energy-efficient home retrofits are a critical element of climate change mitigation strategies. For communities already producing electricity from renewable sources, like Juneau, incentivizing ductless heat pump upgrades and energy efficiency retrofits provides an accessible pathway to transition households away from oil heating and drastically reduce emissions. Thermalize Juneau is a pilot program that will explore a repeatable framework for accelerating the adoption of energy-efficient technology in Alaska communities. This poster will describe Thermalize Juneau's goals and framework, progress through spring 2021, and future plans.

ductless heat pumps↗

Acute wood smoke exposure is associated with cell-specific hippocampal transcriptomic responses in an accelerated ovarian failure mouse model

Background Wildfire events are increasing in frequency and intensity, and aging individuals demonstrate heightened biological susceptibility to air pollution exposures including increased risk of neurological sequelae. Declining ovarian hormones levels that occur with aging in females along with associated systemic physiological and inflammatory changes may contribute to increased cerebral vulnerability to air pollution, representing a potential but underexplored mechanism. Menopause and the menopausal transition represent a period of profound physiological change that affects cardiovascular, neurological, and immune health. Methods We tested whether peri-menopausal–like hormonal status amplifies hippocampal responses to acute wood smoke (WS) using an ovary-intact, 4-vinylcyclohexene diepoxide (VCD) model of moderate accelerated ovarian failure (AOF) in female C57BL/6 mice. Animals were exposed to HEPA-filtered air (FA) or WS for 4 h/day over 2 consecutive days (∼0.5 mg/m³). Exposure characterization confirmed a complex mixture of combustion products with significant levels of both trace metals and gas release during WS exposure. Results Spatial transcriptomics (10x Visium; n = 4 sections/group) with automated cell-type annotation identified astrocytes, GABAergic and glutamatergic neurons, oligodendrocytes, revealed cell type-specific transcriptional alterations following WS exposure. Distinct transcriptional patterns were observed across all identified neuronal and glial cell populations. Conclusion Together, these findings define a cell-type specific transcriptomic framework describing how WS exposure and ovarian hormone decline interact to influence hippocampal responses and identify potential cellular pathways relevant to hippocampal vulnerability.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Advanced Modeling of Conventional Particle Accelerators

SciDAC-5 goals: Deliver particle accelerator and beam simulations tools that go beyond the current state of the art, up to the realization of virtual twins of particle accelerators, enabling design and modeling of particle accelerators at unprecedented speed, levels of accuracy, and realism; and apply these tools to key accelerator facilities relevant to DOE HEP (such as PIP-II/DUNE, FACET-II).

43 PARTICLE ACCELERATORS↗

An integrated Imaging and Modeling Toolbox for Accelerated Development of Root-focused Crops at Field Scales

This project, led by Yuxin Wu at the LBNL, was proposed to develop an integrated imaging-modeling toolbox allowing for accelerated development of root-focused crops at field scales. Our approach is based on a novel root phenotyping method we termed Tomographic Electrical Rhizosphere Imaging (TERI), which provides advanced phenotyping of key root traits by sensing the electrical impedance response of roots and soil to external electrical excitations, and mapping these responses to root and soil properties. The role of Noble Research Institute in this project is to assist TERI model establishment and validation by collecting various above- and below-ground traits data from container- and field-grown plants (mainly wheat). Over the last three years, Noble Research Institute has conducted 3 phases of experiments as proposed including: (Phase 1) greenhouse and hoop-house experiments in wheat and pecan plants grown under controlled container-grown conditions; (Phase 2) space-planted individual plant field experiment using different wheat varieties; and (Phase 3) small-plot planted field experiment using different wheat varieties, as well as using different small grains species including wheat, rye, triticale, oat and barley for increasing plant variation.

60 APPLIED LIFE SCIENCES↗

Progress Towards Accelerating the Unified Model on Hybrid Multi-Core Systems

The cloud microphysics scheme, CASIM, and the radiation scheme, SOCRATES, are two computationally intensive parts within the Met Office's Unified Model (UM). This study enables CASIM and SOCRATES to use accelerated multi-core systems for optimal computational performance of the UM. Using profiling to guide our efforts, we refactored the code for optimal threading and kernel arrangement and implemented OpenACC directives manually or through the CLAW source-to-source translator. Initial porting results achieved 10.02x and 9.25x speedup in CASIM and SOCRATES respectively on 1 GPU compared with 1 CPU core. A granular performance analysis of the strategy and bottlenecks are discussed. These improvements will enable UM to run on heterogeneous computers and a path forward for further improvements is provided.

Zhang, Wei↗

Toward a Holistic Performance Evaluation of Large Language Models Across Diverse AI Accelerators

Artificial intelligence (AI) methods have become critical in scientific applications to help accelerate scientific discovery. Large language models (LLMs) are being considered a promising approach to address some challenging problems because of their superior generalization capabilities across domains. The effectiveness of the models and the accuracy of the applications are contingent upon their efficient execution on the underlying hardware infrastructure. Specialized Al accelerator hardware systems have recently become available for accelerating Al applications. However, the comparative performance of these AI accelerators on large language models has not been previously studied. In this paper, we systematically study LLMs on multiple AI accelerators and GPUs and evaluate their performance characteristics for these models. We evaluate these systems with (i) a micro-benchmark using a core transformer block, (ii) a GPT-2 model, and (iii) an 1,I,M-driven science use case, GenSLM. We present our findings and analyses of the models' performance to better understand the intrinsic capabilities of AI accelerators. Furthermore, our analysis takes into account key factors such as sequence lengths, scaling behavior, and sensitivity to gradient accumulation steps.

Emani, Murali↗