Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix multiplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Sparse matrix‐vector and matrix‐multivector products for the truncated SVD on graphics processors

Summary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm.

Aliaga, José I.↗

Developments in Performance and Portability for MadGraph5_aMC@NLO

Event generators simulate particle interactions using Monte Carlo techniques, providing the primary connection between experiment and theory in experimental high energy physics. These software packages, which are the first step in the simulation worflow of collider experiments, represent approximately 5 to 20% of the annual WLCG usage for the ATLAS and CMS experiments. With computing architectures becoming more heterogeneous, it is important to ensure that these key software frameworks can be run on future systems, large and small. In this contribution, recent progress on porting and speeding up the Madgraph5_aMC@NLO event generator on hybrid architectures, i.e. CPU with GPU accelerators, is discussed. The main focus of this work has been in the calculation of scattering amplitudes and "matrix elements", which is the computational bottleneck of an event generation application. For physics processes limited to QCD leading order, the code generation toolkit has been expanded to produce matrix element calculations using C++ vector instructions on CPUs and using CUDA for NVidia GPUs, as well as using Alpaka, Kokkos and SYCL for multiple CPU and GPU architectures. Performance is reported in terms of matrix element calculations per time on NVidia, Intel, and AMD devices. The status and outlook for the integration of this work into a production release usable by the LHC experiments, with the same functionalities and very similar user interfaces as the current Fortran version, is also described.

Valassi, Andrea↗

A New Model for Simulating the Imbibition of a Wetting-Phase Fluid in a Matrix-Fracture Dual Connectivity System

The imbibition experiment is an effective approach for measuring petrophysical properties of porous media, with many such experiments performed over the past decade. Quite some empirical, analytical, and numerical models have been developed to simulate spontaneous imbibition of the wetting phase fluid into porous media, but limitations still exist. In previous studies, the imbibition process has been considered to give a piston-like displacement or the porous medium modeled as multiply-sized pores linked with bonds; both approaches fail to yield comprehensive results due to their neglect of the presence of irregular fractures or nonuniform flow paths through the matrix. By building a numerical model for simulating laboratory-scale experimental data, we performed imbibition tests on several fractured Barnett Shale samples having fractures either parallel ( P ) or transverse ( T ) to the bedding plane and used MATLAB to build a new numerical model by combining the imbibition process in fractures and the matrix using concepts from percolation theory. The experimental data show that the rocks with P -direction fractures have a more steady increase of imbibition rates than the case of T -direction one. As the shale matrix with low pore connectivity hampers the upward water movement, the imbibition rate of shales with T -direction fractures will decrease suddenly after the bottom layer in contact with water is saturated during the initial period. This wetting phase movement (WPM) model can simulate 3D porous media with 2D fractures. The rate of imbibition by fractured porous media is associated with physical parameters such as porosity and fracture distribution (e.g., the number and angle of fractures). Using Monte Carlo methods, we examined fracture parameters and predicted elapsed time and cumulative water imbibition, for the Barnett Shale samples. The results show that the rate of imbibed water mass is sensitive to the number of fractures directly connected to water source, and the connectivity between two neighboring grid cells is a key parameter for the wetting-front progression. The findings of this study can help to better understand the imbibition process with multiple influencing processes and factors in fractured-matrix rocks. Although the experiments, data simulation, and prediction results are based only on Barnett Shale samples, the model is readily applicable to imbibition tests of other fractured rocks to show the spatial and temporal behavior during a dynamic imbibition process that are not easily captured experimentally.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A quantitative phase-field model for gas bubble evolution in UO2

Due to the large formation energy of vacancies and noble gas atoms in the form of interstitials or substitutional atoms in nuclear fuel (UO2), the thermodynamic equilibrium concentrations of these species are very low in the nuclear fuel matrix even at very high temperatures, which imposes difficulties upon the quantitative study of bubble evolution via the phase-field method. In this study, a quantitative phase-field model is proposed to deal with this problem. The system’s free energy density is derived according to the principles of thermodynamics, with consideration of the elastic effect and with the use of real material parameters from experiments. The model is useful for the study of the kinetics of gas bubble growth with very dilute concentrations of vacancy and gas atoms in the matrix. This model is applied to study single bubble growth and multiple bubble growth under various concentrations of vacancy and gas atoms and at various temperatures. The elastic effect and the effects of the generation rate of vacancies and gas atoms on bubble growth are analyzed.

low defect concentration, nuclear fuel, gas bubble↗

Performance of modern color decompositions for standard candle LHC tree amplitudes

In the last decade, developments of matrix element and phase space generators have focused on providing good efficiency and maximal flexibility and automation for a wide range of physical processes. However, as recent studies have shown, they are a major bottleneck in the established Monte Carlo event generator toolchains. With the advent of the HL-LHC and ever rising precision requirements, future developments will need to focus on computational performance, especially at intermediate to large jet multiplicities. We present the novel BlockGen family of fast matrix element algorithms that are amenable for GPU acceleration, making use of modern, minimal color decompositions. Moreover, we discuss the performance achieved for standard candle processes such as V +jets and tt̄+jets production.

Bothmann, E. [Gottingen U.]↗

Crystallography of Reactive Intermediates

Metal–ligand (M–L) multiply bonded complexes hold a central place in inorganic chemistry and catalysis: From fundamental and historical perspectives, these species have played a critical role in the articulation of important bonding principles (i.e. the vanadyl ion in the development of molecular orbital theory); from a practical perspective, these species are critical intermediates in a variety of chemical reactions (i.e. N 2 and O 2 reduction, H 2 O oxidation, and C–H functionalization). For many high-valent, early metal complexes, overlap of ligand-based electrons with vacant π-symmetry orbitals leads to strong M–L multiple bonds. The stability of these species enables straightforward characterization with a suite of standard spectroscopic and diffraction-based experiments. Mid- and late-transition metal complexes, with attendant higher d-electron counts, often support more reactive M–L multiply bonded fragments. The reactivity of these species simultaneously renders them attractive intermediates for catalysis but challenging synthetic targets to observe and characterize. A number of important strategies have been advanced to enable experimental characterization of mid- to late-metal–ligand multiply bonded species. Synthetic manipulation of the coordination geometry and ligand donicity, as well as introduction of sterically encumbering ligands, have each emerged as powerful methods to tame the inherent reactivity of kinetically labile M–L multiple bonds. While these efforts have resulted in families of well-characterized complexes and provided critical insights regarding structure and bonding, the synthetic derivatization required to stabilize M–L fragments of interest often obviates the substrate functionalization activity relevant to catalysis. Photochemical synthesis of reactive species provides a conceptually attractive strategy to generate reactive M–L fragments under conditions compatible with time-resolved or cryogenic steady-state characterization, and photogeneration has enabled observation of a number of reactive M–L fragments. The suite of tools available to characterize photogenerated reactive species is often more limited than typical for kinetically stabilized complexes and structural characterization is typically not possible. Recently, photocrystallographic experiments, in which reactive M–L multiply bonded intermediates are generated within single-crystal matrices, have been advanced as a strategy to interrogate the structures of reactive intermediates in C–H functionalization. Lastly, this Comment describes the historical antecedents to these experiments, highlights examples of photocrystallographic characterization of reactive intermediates, and discusses future opportunities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A System-Level Cost Modeling Framework for Design for Remanufacturing: A Case Study of an Agricultural Machine Transmission

Remanufacturing offers significant environmental and economic benefits by restoring end-of-life products to as-new conditions. Although extensive research has been conducted on the topic, the adoption of remanufacturing practices remains limited across various industries. A primary barrier to broader implementation is the substantial upfront investment required, which necessitates reliable cost modeling to justify potential future savings. Most existing models treat components independently and ignore inter-component dependencies. We develop a probabilistic, system-level cost modeling framework that integrates reliability, reusability, and a dependency matrix to capture cascading effects across components over multiple life cycles. Our model identifies those critical components that maximize remanufacturing benefits across a product's many lives. A toy example and an industry case study (John Deere PowrQuad transmission subassembly) illustrate how design alternatives affect cumulative cost. Using a Monte Carlo simulation (MCS) to perform life cycle cost analysis on the system with different design changes, we show the normalized average cost savings after three remanufacturing cycles. Accounting for dependencies meaningfully alters cost projections and ignoring them underestimates accumulated cost by up to 20% in our examples. Furthermore, the results of our study confirm that accounting for component interdependencies is necessary to produce cost estimates that meet industry standards.

Life Cycle Analysis and Design↗

Time-resolved tracking of cellulose biosynthesis and assembly during cell wall regeneration in live Arabidopsis protoplasts

Cellulose, the most abundant polysaccharide on earth composing plant cell walls, is synthesized by coordinated action of multiple enzymes in cellulose synthase complexes embedded within the plasma membrane. Multiple chains of cellulose fibrils form intertwined extracellular matrix networks. It remains largely unknown how newly synthesized cellulose is assembled into an intricate fibril network on cell surfaces. Here, we have established an in vivo time-resolved imaging platform to continuously visualize cellulose biosynthesis and fibril network assembly onArabidopsis thalianaprotoplast surfaces as the primary cell wall regenerates. Our observations provide the basis for a model of cellulose fibril network development in protoplasts driven by an interplay of multiscale dynamics that includes rapid diffusion and coalescence of nascent cellulose fibrils, processive elongation of single fibrils, and cellulose fibrillar network rearrangement during maturation. This study provides fresh insights into the dynamic and mechanistic aspects of cell wall synthesis at the single-cell level.

Science & Technology - Other Topics↗

Nitrogen Evaluation [Slides]

This presentation includes a R-Matrix Analysis of 15N system with SAMMY, Discussion of Boundary Conditions in R-Matrix Analyses and a discussion of Analyses with Multiple Incident Channels [Slides]

15-N↗

Nuclear Criticality Safety Integral Experiment Covariance Determination

Integral benchmarks for criticality safety and nuclear data validation require expensive uncertainty quantification studies. Commonly, the uncertainty quantification ignores correlations between experiments that share components. Experiments such as the TEX (Thermal/Epithermal eXperiments) campaigns consist of many shared parts between experiments, such as fuel, which creates a strong correlation in their errors. While these correlations are known to exist, they are often not estimated due to the complexity of such calculations. This paper describes a software package that uses an intuitive method of determining the covariance for each of the experimental components, providing a correlation matrix for each family of parts across the multiple cases examined within a benchmark. The code uses the TEX-HEU campaign as a proof of concept, and we show that the correlations can be calculated with information commonly found in ICSBEP (International Criticality Safety Benchmark Evaluation Project) benchmarks. The estimated covariances are used in χ 2 trending studies to evaluate their impact on nuclear data validation. The covariance determination code can be easily integrated into current benchmark evaluations as well as reevaluating legacy benchmark uncertainties. Uncertainty correlation calculations should become the baseline for criticality safety integral experiment benchmarks and can now be easily calculated with the described software package.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Nuclear Criticality Safety Integral Experiment Covariance Determination

Integral benchmarks for criticality safety and nuclear data validation require expensive uncertainty quantification studies. Commonly, the uncertainty quantification ignores correlations between experiments that share components. Experiments such as the TEX (Thermal/Epithermal eXperiments) campaigns consist of many shared parts, such as fuel, which create a strong correlation in their uncertainties. While these correlations are known to exist, they are often not estimated due to the complexity of such calculations. This paper describes a software package that uses an intuitive method of determining the covariance for each of the experimental components, providing a correlation matrix for each family of parts across the multiple cases examined within a benchmark. The code uses the TEX-HEU campaign as a proof of concept, and we show that the correlations can be calculated with information commonly found in ICSBEP (International Criticality Safety Benchmark Evaluation Project) benchmarks. The estimated covariances are used in χ 2 trending studies to evaluate their impact on nuclear data validation. Without covariances, χ 2 per degree of freedom was calculated as 2.203 and with covariances it was 1.179. The difference shows that omitting covariance information may cause overly pessimistic bias quantifications. The covariance determination code can be easily integrated into current benchmark evaluations as well as reevaluating legacy benchmark uncertainties. Uncertainty correlation calculations should become the baseline for criticality safety integral experiment benchmarks and can now be easily calculated with the described software package.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

At-power subcritical multiplication in the Advanced Test Reactor during nuclear requalification testing

Power division information during nuclear requalification of the Advanced Test Reactor (ATR) is of considerable interest as an importance function for observed changes to core reactivity. The degree to which a given physical subdivision of a critical reactor acts as a neutron source for other lobes is not analytically characterized for general application. When ATR operates at power, individual power-producing lobes rely on each other as neutron sources in order to maintain constant power, which in general requires either exactly critical multiplication within a reactor or an external neutron source. Here, this work shows that fuel element and lobe powers in ATR can be related with subcritical multiplication theory. Subcritical multiplication factors are computed with a physically validated analytical method based on actual at-power operation, quantifying for each lobe its dependence on other lobes as an external neutron source. This explanation is significant for ATR due to the desire to irradiate a large variety of experiments simultaneously, each having its impact on the core neutron population. For any physical subdivision of any other critical reactor, it is likewise true that the subdivision undergoes only subcritical multiplication.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

DXRD : a user-friendly suite of two- and multiple-beam dynamical X-ray diffraction programs

The DXRD program suite consisting of a series of dynamical theory programs is introduced for computing dynamical X-ray diffraction from single crystals. Its interactive graphical user interfaces (GUIs) allow general users to make complicated calculations with minimal effort. It can calculate plane-wave Darwin curves of single crystals (or multiple crystals) for both the Bragg and Laue cases, including grazing-incidence diffraction and backward diffraction (with Bragg angles approaching 90°). It is also capable of simulating rocking curves for divergent incident X-ray beams with finite bandwidths. A unique feature of DXRD is that it provides a convenient GUI-based multiple-beam diffraction program that can accurately compute arbitrary N-beam diffraction of any geometry using a universal 4N × 4N matrix method. DXRD also provides a mapping program for plotting all the multiple-beam diffraction lines (monochromator glitches) in the azimuth–energy coordinate system. All these functions make DXRD a convenient and powerful software tool for designing crystal-based synchrotron/X-ray optics (monochromators, analyzers, polarizers, phase plates etc.) and for crystal characterization, X-ray spectroscopy and X-ray diffraction teaching.

Bragg reflection↗

In Situ Oxygen Isotope Determination in Serpentine Minerals by SIMS: Addressing Matrix Effects and Providing New Insights on Serpentinisation at Hole BA1B (Samail ophiolite, Oman)

The ability to constrain the petrogenesis of multiple serpentine generations recorded at the microscale is crucial for estimating the extent and conditions of modern versus fossil serpentinisation in ophiolites. To address matrix bias effects during oxygen isotope analysis by SIMS, we present the first investigation analysing antigorite in the compositional range Mg# = 77.5–99.5 mole %, using a CAMECA IMS–1280 secondary ion mass spectrometer. Spot–to–spot homogeneity is ≤ 0.5‰ (2s) for the new antigorite reference materials. The relative bias between antigorite reference materials with different Mg/Fe ratios is described by a second order polynomial, and a maximum difference in bias of ~ 1.8‰ was measured for Mg# ~ 78 to 100. We observed a bias up to ~ 1.0‰ between lizardite and antigorite attributed to their different crystal structure. Orientation effects up to ~ 1‰ were observed in chrysotile. The new analytical protocol allowed the identification of oxygen isotope zoning up to ~ 7‰ in serpentine minerals from two serpentinites recovered from an area of active serpentinisation in the Samail ophiolite. Furthermore in situ analysis is capable of resolving isotopic heterogeneity that may directly reflect changes in the physical and chemical conditions of multiple serpentinisation events in the Samail ophiolite.

47 OTHER INSTRUMENTATION↗

Systematic planning of moving target defence for maximising detection effectiveness against false data injection attacks in smart grid

Abstract Moving target defence (MTD) has been gaining traction to thwart false data injection attacks against state estimation (SE) in the power grid. MTD actively perturbs the reactance of transmission lines equipped with distributed flexible AC transmission system (D‐FACTS) devices to falsify the attacker's knowledge about the system configuration. However, the existing literature has not systematically studied what influences the detection effectiveness of MTD and how it can be improved based on the topology analysis. These problems are tackled here from the perspective of an MTD plan in which the D‐FACTS placement is determined. We first exploit the relation between the rank of the composite matrix and the detecting effectiveness. Then, we rigorously derive upper and lower bounds on the attack detecting probability of MTDs with a given rank of the composite matrix. Furthermore, we analyse existing planning methods and highlight the importance of bus coverage by D‐FACTS devices. To improve the detection effectiveness, we propose a novel graph theory–based planning algorithm to retain the maximum rank of the composite matrix while covering all necessary buses. Comparative results on multiple systems show the high detecting effectiveness of the proposed algorithm in both DC‐ and AC‐SE.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Isospin 0 and 2 two-pion scattering at physical pion mass using all-to-all propagators with periodic boundary conditions in lattice QCD

A study of two-pion scattering for the isospin channels, 𝐼 = 0 and 𝐼 = 2, using lattice QCD is presented. Möbius domain-wall fermions, on top of the Iwasaki-DSDR gauge action for gluons with periodic boundary conditions, are used for the lattice computations, which are carried out on two ensembles of gauge field configurations generated by the RBC and UKQCD Collaborations with physical masses, inverse lattice spacings of 1.023 and 1.378 GeV, and spatial extents of 𝐿 = 4.63 and 4.58 fm, respectively. The all-to-all propagator method is employed to compute a matrix of correlation functions of two-pion operators. The generalized eigenvalue problem (GEVP) is solved for a matrix of correlation functions to extract phase shifts with multiple states—two pions with a nonzero relative momentum, as well as two pions at rest. Our results for phase shifts for both the 𝐼 = 0 and 𝐼 = 2 channels are consistent with the Roy equation and chiral perturbation theory, though at this preliminary stage our errors for 𝐼 = 0 are large. An important outcome of this work is that we are successful in extracting two-pion excited states, which are useful for studying 𝐾 → 𝜋⁢𝜋 decay, on physical-mass ensembles using the GEVP.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗