Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “dot product”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Mode-multiplexed photonic integrated vector dot-product core from inverse design

Photonic computing has the potential to harness the full degrees of freedom (DOFs) of the light field, including the wavelength, spatial mode, spatial location, phase quadrature, and polarization, to achieve a higher level of computing parallelism and scalability than digital electronic processors. While multiplexing using the wavelength and other DOFs can be readily integrated on silicon photonics platforms with compact footprints, conventional mode-division multiplexed (MDM) photonic designs occupy areas exceeding tens to hundreds of microns for a few spatial modes, significantly limiting their scalability. Here, we utilize inverse design to demonstrate an ultracompact photonic computing core that calculates vector dot products based on MDM coherent mixing. Our dot-product core integrates the functionalities of two-mode multiplexers and one multimode coherent mixer within a nominal footprint of 5 μm x 3 μm . We have experimentally demonstrated computing examples on the fabricated dot-product core, including complex number multiplication and motion estimation using optical flow. The compact dot-product core design enables large-scale on-chip integration in a parallel photonic computing primitive cluster for high-throughput scientific computing and computer vision tasks.

97 MATHEMATICS AND COMPUTING↗

A Framework for Error-Bounded Approximate Computing, with an Application to Dot Products

Approximate computing techniques, which trade off the computation accuracy of an algorithm for better performance and energy efficiency, have been successful in reducing computation and power costs in several domains. However, error sensitive applications in high-performance computing are unable to benefit from existing approximate computing strategies that are not developed with guaranteed error bounds. While approximate computing techniques can be developed for individual high-performance computing applications by domain specialists, this often requires additional theoretical analysis and potentially extensive software modification. Hence, the development of low-level error-bounded approximate computing strategies that can be introduced into any high-performance computing application without requiring additional analysis or significant software alterations is desirable. In this paper, we provide a contribution in this direction by proposing a general framework for designing error-bounded approximate computing strategies and apply it to the dot product kernel to develop \bf qdot---an error-bounded approximate dot product kernel. Following the introduction of qdot, here we perform a theoretical analysis that yields a deterministic bound on the relative approximation error introduced by qdot. Empirical tests are performed to illustrate the tightness of the derived error bound and to demonstrate the effectiveness of qdot on a synthetic dataset, as well as two scientific benchmarks---the conjugate gradient (CG) and power methods. In some instances, using qdot for the dot products in CG can result in many components being quantized to half precision without increasing the iteration count required for convergence to the same solution as CG using a double precision dot product.

97 MATHEMATICS AND COMPUTING↗

sKokkos: Enabling Kokkos with Transparent Device Selection on Heterogeneous Systems using OpenACC

This paper presents a new feature to enable Kokkos with transparent device selection. For application developers, it is not easy toidentify which device is the most appropriate to use in a heterogeneous system, since this depends on the characteristics of both the application and the hardware. In Kokkos, a backend is associated with one specific programming model/hardware. Programmers decide which backend to use at compilation time. This new feature implemented on the OpenACC backend eliminates the burden of deciding which device to use, providing a highly productive programming solution for Kokkos applications. This work includes implementation details and a performance study conducted with a set of mini-benchmarks (i.e., AXPY and dot product), kernels (Lattice-Bolzmann method), and two mini-apps (LULESH and miniFE) on two heterogeneous systems with different hardware capabilities. This new Kokkos feature provides high accelerations of up to 35× thanks to automatic and transparent device selection.

Lee, Seyong↗

KokkACC: Enhancing Kokkos with OpenACC

Template metaprogramming is gaining popularity as a high-level solution for achieving performance portability on heterogeneous computing resources. Kokkos is a representative approach that offers programmers high-level abstractions for generic programming while most of the device-specific code generation and optimizations are delegated to the compiler through template specializations. For this, Kokkos provides a set of device-specific code specializations in multiple back ends, such as CUDA and HIP. Unlike CUDA or HIP, OpenACC is a high-level and directive-based programming model. This descriptive model allows developers to insert hints (pragmas) into their code that help the compiler to parallelize the code. The compiler is responsible for the transformation of the code, which is completely transparent to the programmer. This paper presents an OpenACC back end for Kokkos: KokkACC. As an alternative to Kokkos’s existing device-specific back ends, KokkACC is a multi-architecture back end providing a high-productivity programming environment enabled by OpenACC’s high-level and descriptive programming model. Moreover, we have observed competitive performance; in some cases, KokkACC is faster (up to 9×) than NVIDIA’s CUDA back end and much faster than OpenMP’s GPU offloading back end. This work also includes implementation details and a detailed performance study conducted with a set of mini-benchmarks (AXPY and DOT product) and three mini-apps (LULESH, miniFE and SNAP, a LAMMPS proxy mini-app).

Valero Lara, Pedro↗

Toward efficient polynomial preconditioning for GMRES

Here, we present a polynomial preconditioner for solving large systems of linear equations. The polynomial is derived from the minimum residual polynomial (the GMRES polynomial) and is more straightforward to compute and implement than many previous polynomial preconditioners. Our current implementation of this polynomial using its roots is naturally more stable than previous methods of computing the same polynomial. We implement further stability control using added roots, and this allows for high degree polynomials. We discuss the effectiveness and challenges of root-adding and give an additional check for stability. In this article, we study the polynomial preconditioner applied to GMRES; however it could be used with any Krylov solver. This polynomial preconditioning algorithm can dramatically improve convergence for some problems, especially for difficult problems, and can reduce dot products by an even greater margin.

97 MATHEMATICS AND COMPUTING↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Novel Geometric Operations for Linear Programming

This report summarizes the work performed under the project "Linear Programming in Strongly Polynomial Time." Linear programming (LP) is a classic combinatorial optimization problem heavily used directly and as an enabling subroutine in integer programming (IP). Specifically IP is the same as LP except that some solution variables must take integer values (e.g. to represent yes/no decisions). Together LP and IP have many applications in resource allocation including general logistics, and infrastructure design and vulnerability analysis. The project was motivated by the PI's recent success developing methods to efficiently sample Voronoi vertices (essentially finding nearest neighbors in high-dimensional point sets) in arbitrary dimension. His method seems applicable to exploring the high-dimensional convex feasible space of an LP problem. Although the project did not provably find a strongly-polynomial algorithm, it explored multiple algorithm classes. The new medial simplex algorithms may still lead to solvers with improved provable complexity. We describe medial simplex algorithms and some relevant structural/complexity results. We also designed a novel parallel LP algorithm based on our geometric insights and implemented it in the Spoke-LP code. A major part of the computational step is many independent vector dot products. Our parallel algorithm distributes the problem constraints across processors. Current commercial and high-quality free LP solvers require all problem details to fit onto a single processor or multicore. Our new algorithm might enable the solution of problems too large for any current LP solvers. We describe our new algorithm, give preliminary proof-of-concept experiments, and describe a new generator for arbitrarily large LP instances.

97 MATHEMATICS AND COMPUTING↗

Vector-Matrix Multiplication Engine for Neuromorphic Computation with a CBRAM Crossbar Array [Slides]

The core function of many neural network algorithms is the dot product, or vector matrix multiply (VMM) operation. Crossbar arrays utilizing resistive memory elements can reduce computational energy in neural algorithms by up to five orders of magnitude compared to conventional CPUs. Moving data between a processor, SRAM, and DRAM dominates energy consumption. By utilizing analog operations to reduce data movement, resistive memory crossbars can enable processing of large amounts of data at lower energy than conventional memory architectures.

97 MATHEMATICS AND COMPUTING↗

Light-driven hydrogen production with CdSe quantum dots and a cobalt glutathione catalyst

A photocatalytic hydrogen (H 2 ) production system is reported using glutathione (GSH)-capped CdSe QDs with a cobalt precatalyst, yielding 130 000 mol H 2 per mol cobalt over 48 hours. Analysis of the reaction mixtures after catalysis indicates that the active catalyst is a labile complex of cobalt and GSH formed in situ.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Linking emergent phenomena and broken symmetries through one-dimensional objects and their dot/cross products

Abstract The symmetry of the whole experimental setups, including specific sample environments and measurables, can be compared with that of specimens for observable physical phenomena. We, first, focus on one-dimensional (1D) experimental setups, independent from any spatial rotation around one direction, and show that eight kinds of 1D objects (four; vector-like, the other four; director-like), defined in terms of symmetry, and their dot and cross products are an effective way for the symmetry consideration. The dot products form a Z 2 × Z 2 × Z 2 group with Abelian additive operation, and the cross products form a Z 2 × Z 2 group with Abelian additive operation or Q 8 , a non-Abelian group of order eight, depending on their signs. Those 1D objects are associated with characteristic physical phenomena. When a 3D specimen has symmetry operational similarity (SOS) with (identical or lower, but not higher, symmetries than) an 1D object with a particular phenomenon, the 3D specimen can exhibit the phenomenon. This SOS approach can be a transformative and unconventional avenue for symmetry-guided materials designs and discoveries.

Physics↗

Non-equilibrium entropy production and information dissipation in a non-Markovian quantum dot

This study measures trajectory-level entropy production and information dissipation in a driven, non-Markovian quantum dot using time-resolved optical dynamics and machine-learning-based analysis. Although not a 2D-material system, it is relevant because it demonstrates quantitative extraction of nonequilibrium dynamics from nanoscale optical fluctuations, which is conceptually connected to the proposed studies of transient charge and spin dynamics at interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Environmentally Friendly Production of High-Quality and Multifunctional Carbon Quantum Dots from Coal

Researchers from the University of Wyoming and the University of Utah worked jointly to study the valorization of coal to carbon quantum dots (CQDs), a value-added product with a broad spectrum of applications. The CQDs were produced by an environmentally facile hydrothermal method, and the experimental factors influencing the properties of CQDs were investigated. We subsequently explored the applications of CQDs as co-sensitizers of dye-sensitized solar cells (DSSC) and photocatalysts of water treatment. As an outlook, techno-economic and environmental analysis studied the feasibility of mass production of CQDs.

01 COAL, LIGNITE, AND PEAT↗

Trace level of atomic copper in N-doped graphene quantum dots switching the selectivity from C 1 to C 2 products in CO electroreduction

To unravel the relationship between trace Cu on the metal-free catalysts toward CO/CO 2 reduction reaction (CO/CO 2 RR), we investigated the effect of trace Cu loading in N-doped graphene quantum dots (NGQDs) on CO/CO 2 RR. A general trend is that increasing the Cu loading in NGQDs switches the selectivity from C 1 (CH 4 ) to C 2 products in CORR. When 2.5 μg/cm 2 Cu with the atomic size is loaded on NGQDs, the selectivity shifts from 62% Faradaic efficiency (FE) of CH 4 to 52% FE of C 2 products in CORR. Further increasing the atomic Cu loading to 3.8 μg/cm 2 promotes the FE of C 2 products to 78%. CO 2 RR requires one order of magnitude higher Cu loading than CORR to switch the selectivity from C 1 to C 2 products due to the low partial pressure of CO. Finally, this study clarifies the distinct impact of trace (ppm level) Cu on the activity/selectivity between CORR and CO 2 RR.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Alternate InP synthesis with aminophosphines: solution–liquid–solid nanowire growth

Indium phosphide nanowires are important components in high-speed electronics and optoelectronics, including photodetectors and photovoltaics. However, most syntheses either use high-temperature and costly vapor-phase methodology or highly toxic and pyrophoric tris(trimethylsilyl)phosphine. To expand on the success of the aminophosphine-based InP colloidal quantum dot synthesis, we developed a synthesis for thin (~11 nm) zinc blende InP nanowires at 180 °C using indium tris(trifluoroacetate) and tris(diethylamino)phosphine. A flat nanoribbon morphology was identified by transmission electron and atomic force microscopy analysis, with the stoichiometric (110) lattice plane exposed. Nanowire growth proceeded through a solution–liquid–solid mechanism from in situ-formed indium metal nanoparticles. Molecular byproducts of tris(oleylamino)phosphine oxide and N-oleyltrifluoroacetamide observed by 31 P and 19 F NMR spectroscopy inform a proposed mechanism of indium reduction by the aminophosphine. Morphological control over the nanowire product was achieved by varying the phosphorus injection to control the aspect ratio, the In : P ratio to toggle between nanowires and multipods, and the pre-hot injection evacuation step to favor a quantum dot product. Furthermore, replacing the indium precursor with indium tris(trifluoromethanesulfonate) was found to make bulk zinc blende InP nanowires with an average diameter of >250 nm and tens of microns in length.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Influence of QD photosensitizers in the photocatalytic production of hydrogen with biomimetic [FeFe]-hydrogenase. Comparative performance of CdSe and CdTe

Photocatalytic systems comprising a hydrogenase-type catalyst and CdX (X = S, Se, Te) chalcogenide quantum dot (QD) photosensitizers show extraordinary hydrogen production rates under visible light excitation. What remains unknown is the mechanism of energy conversion in these systems. In this work, we have explored this question by comparing the performance of two QD sensitizers, CdSe and CdTe, in photocatalytic systems featuring aqueous suspensions of a [Fe 2 (μ-1,2-benzenedithiolate) CO 6 ] catalyst and an ascorbic acid sacrificial agent. Overall, the hydrogen production yield for CdSe-sensitized reactions QDs was found to be 13 times greater than that of CdTe counterparts. According to emission quenching experiments, an enhanced performance of CdSe sensitizers reflected a greater rate of electron transfer from the ascorbic acid (k Asc ). The observed difference in the QD-ascorbic acid charge transfer rates between the two QD materials was consistent with respective driving forces for these systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gradient synthesis of carbon quantum dots and activated carbon from pulp black liquor for photocatalytic hydrogen evolution and supercapacitor

Black liquor (BL) is a by-product of the chemical pulping industry and is mainly used as a low-value fuel; however, its potential to produce high-value products has not been fully exploited. In this study, a green and simple strategy is reported for the gradient production of Na + -functionalized carbon quantum dots (Na + -CQDs) for the first time, N and S co-doped CQDs (N/S-CQDs), and N and S co-doped KOH-activated carbon (N/S-KAC) from BL by dialysis, hydrothermal carbonization and activation-carbonization, respectively. Due to the good electron trapping ability, photoluminescence and promising up-conversion luminescence of CQDs, the hydrogen evolution efficiency of Na + -CQDs/TiO 2 and N/S-CQDs/TiO 2 photocatalysts was improved by 2.45 and 1.46 times, respectively, compared with pure TiO 2 . N/S-KAC with a high specific surface area of 2294 m 2 g -1 provides an excellent specific capacitance of 253 F g -1 at 0.5 A g -1 and a promising energy density of 26.92 Wh kg -1 under a power density of 566 W kg -1 for the fabricated symmetrical supercapacitor. Moreover, the electrode material has good cycling stability with a capacitance retention of ~ 93.91% after 5000 cycles. In conclusion, this pathway provides a versatile and scalable approach for the construction and co-production of nanostructured materials, photocatalysts and energy storage devices.

36 MATERIALS SCIENCE↗