Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Moment-based adaptive time integration for thermal radiation transport

Here, in this paper we develop a framework for moment-based adaptive time integration of deterministic multifrequency thermal radiation transpot (TRT). We generalize our recent semi-implicit-explicit (IMEX) integration framework for gray TRT to multifrequency TRT, and also introduce a semi-implicit variation that facilitates higher-order integration of TRT, where each stage is implicit in all components except opacities. To appeal to the broad literature on adaptivity with Runge–Kutta methods, we derive new embedded methods for four asymptotic preserving IMEX Runge–Kutta schemes we have found to be robust in our previous work on TRT and radiation hydrodynamics. We then use a moment-based high-order-low-order representation of the transport equations. Due to the high dimensionality, memory is always a concern in simulating TRT. We form error estimates and adaptivity in time purely based on temperature and radiation energy, for a trivial overhead in computational cost and memory usage compared with the base second order integrators. We then test the adaptivity in time on the tophat and Larsen problem, demonstrating the ability of the adaptive algorithm to naturally vary the timestep across 4–5 orders of magnitude, ranging from the dynamical timescales of the streaming regime to the thick diffusion limit.

97 MATHEMATICS AND COMPUTING↗

Probing boron vacancy defects in hBN via single spin relaxometry

Spin defects in solids offer promising platforms for quantum sensing and memory due to their long coherence times and optical addressability. Here, we integrate a single nitrogen-vacancy (NV) center in diamond with scanning probe microscopy to detect, read out, and spatially map spin-based quantum sensors at the nanoscale. Using the boron vacancy ($V$$^{–}_{B}$) center in hexagonal boron nitride—an emerging two-dimensional spin system—as a model, we detect its electron spin resonance indirectly via changes in the spin relaxation time (T 1 ) of a nearby NV center, eliminating the need for optical excitation or fluorescence detection of the $V$$^{–}_{B}$. Cross-relaxation between NV and $V$$^{–}_{B}$ ensembles significantly reduces NV T1, enabling quantitative nanoscale mapping of defect densities beyond the optical diffraction limit and clear resolution of hyperfine splitting in isotopically enriched h 10 B 15 N. Our method demonstrates interactions between spin sensors in 3D and 2D materials, establishing NV centers as versatile probes for characterizing otherwise inaccessible spin defects.

Quantum metrology↗

Reconstructing Hanford worker external doses from photons for epidemiology

The accurate reconstruction of external photon doses is essential for credible radiation epidemiology. This article presents the methodology used to derive dose estimates for 37 012 Hanford Site workers included in the Million Person Study. The approach employs historical dose records from the Hanford Radiation Exposure database and a previous epidemiology study. Bias correction factors specific to dosimeter type and period of use were applied and missing annual doses were estimated using a hierarchical nearby method to estimate deep dose equivalent for each worker. For early years with limited detection sensitivity, missed doses were quantified based on expected time-period-specific, low-dose statistical distributions. The revised dose estimates resulted in lower median and mean career doses than unadjusted data, while increasing the number of person-years with nonzero dose. Sensitivity analyses assessed the influence of bias in dosimetry measurements, missed doses and gap years on dose estimates. Differences in cumulative dose estimates between unadjusted and revised annual estimates are most prominent in the early operational years due to the highest bias during that time period.

dose reconstruction↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sulfurization Engineering of One–Step Low–Temperature MoS 2 and WS 2 Thin Films for Memristor Device Applications

2D materials have been of considerable interest as new materials for device applications. Non-volatile resistive switching applications of MoS 2 and WS 2 have been previously demonstrated; however, these applications are dramatically limited by high temperatures and extended times needed for the large-area synthesis of 2D materials on crystalline substrates. The experimental results demonstrate a one-step sulfurization method to synthesize MoS 2 and WS 2 at 550 °C in 15 min on sapphire wafers. Furthermore, a large area transfer of the synthesized thin films to SiO 2 /Si substrates is achieved. Following this, MoS 2 and WS 2 memristors are fabricated that exhibit stable non-volatile switching and a satisfactory large on/off current ratio (10 3 –10 5 ) with good uniformity. Tuning the sulfurization parameters (temperature and metal precursor thickness) is found to be a straightforward and effective strategy to improve the performance of the memristors. Furthermore, the demonstration of large-scale MoS 2 and WS 2 memristors with a one-step low-temperature sulfurization method with simple strategy to tuning can lead to potential applications such as flexible memory and neuromorphic computing.

36 MATERIALS SCIENCE↗

Traversing large graphs on GPUs with unified memory

Due to the limited capacity of GPU memory, the majority of prior work on graph applications on GPUs has been restricted to graphs of modest sizes that fit in memory. Recent hardware and software advances make it possible to address much larger host memory transparently as a part of a feature known as unified virtual memory. While accessing host memory over an interconnect is understandably slower, the problem space has not been sufficiently explored in the context of a challenging workload with low computational intensity and an irregular data access pattern such as graph traversal. We analyse the performance of breadth first search (BFS) for several large graphs in the context of unified memory and identify the key factors that contribute to slowdowns. Next, we propose a lightweight offline graph reordering algorithm, HALO (Harmonic Locality Ordering), that can be used as a pre-processing step for static graphs. HALO yields speedups of 1.5x-1.9x over baseline in subsequent traversals. Our method specifically aims to cover large directed real world graphs in addition to undirected graphs whereas prior methods only account for the latter. Additionally, we demonstrate ties between the locality ordering problem and graph compression and show that prior methods from graph compression such as recursive graph bisection can be suitably adapted to this problem.

Gera, Prasun↗

Estimating Flexibility Envelopes for Residential Customers From Utility Smart Meter Data: Preprint

Demand response from residential customers has significant potential to support power system operations, but accurate flexibility estimation is challenging due to the limited resolution of advanced metering infrastructure (AMI) data. Most utility AMI measurements are recorded at hourly intervals, with only a small portion at higher resolutions, and even fewer households have appliance-level energy usage data. To address this issue, this paper proposes a two-stage long short-term memory (LSTM) framework for estimating household flexibility envelopes from low-resolution AMI data. In the first stage, the heating, ventilating, and air-conditioning (HVAC) load and non-HVAC loads are estimated by using a model trained on a small set of households with appliance-level profiles. These estimated data are then used to compute the upper- and lower-flexibility bounds, which are subsequently down-sampled to lower-resolution data. In the second stage, these flexibility bounds serve as training inputs for another LSTM model, enabling direct prediction of flexibility envelopes for households with only hourly AMI data. This method is validated using Pecan Street data from two different areas, and the results demonstrate its applicability and effectiveness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Ferroelectricity in Si-Doped Hafnia: Probing Challenges in Absence of Screening Charges

The ability to develop ferroelectric materials using binary oxides is critical to enable novel low-power, high-density non-volatile memory and fast switching logic. The discovery of ferroelectricity in hafnia-based thin films, has focused the hopes of the community on this class of materials to overcome the existing problems of perovskite-based integrated ferroelectrics. However, both the control of ferroelectricity in doped-HfO2 and the direct characterization at the nanoscale of ferroelectric phenomena, are increasingly difficult to achieve. The main limitations are imposed by the inherent intertwining of ferroelectric and dielectric properties, the role of strain, interfaces and electric field-mediated phase, and polarization changes. In this work, using Si-doped HfO2 as a material system, we performed a correlative study with four scanning probe techniques for the local sensing of intrinsic ferroelectricity on the oxide surface. Putting each technique in perspective, we demonstrated that different origins of spatially resolved contrast can be obtained, thus highlighting possible crosstalk not originated by a genuine ferroelectric response. By leveraging the strength of each method, we showed how intrinsic processes in ultrathin dielectrics, i.e., electronic leakage, existence and generation of energy states, charge trapping (de-trapping) phenomena, and electrochemical effects, can influence the sensed response. We then proceeded to initiate hysteresis loops by means of tip-induced spectroscopic cycling (i.e., “wake-up”), thus observing the onset of oxide degradation processes associated with this step. Finally, direct piezoelectric effects were studied using the high pressure resulting from the probe’s confinement, noticing the absence of a net time-invariant piezo-generated charge. Our results are critical in providing a general framework of interpretation for multiple nanoscale processes impacting ferroelectricity in doped-hafnia and strategies for sensing it.

36 MATERIALS SCIENCE↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Disorder, interactions, and their interplay in novel narrow-gap Dirac materials and Weyl semimetals

Progress of the modern day condensed matter physics is to a large extent driven by the synthesis of new materials, advances in their experimental characterization and theoretical description. Recent discoveries of novel gapless Weyl semimetals, such as NaBi,CdAs, and BiTe-based films, in which magnetic dopants essentially suppress the gap, have added to the family of graphene and topological insulators actively investigated over the past decade. With the field of novel semimetals rapidly maturing, its focus necessarily shifts from demonstrations of the feasibility of such materials to their quantitative characterization. While the transport and optical properties of graphene and topological insulators are well captured within the picture of free non-interacting electrons, gapless 3D Weyl semimetals and narrow-gap 2D semiconductors with Dirac spectrum are known to be extremely susceptible to disorder and electron-electron interactions. This susceptibility obscures the manifestations of nontrivial band structure -- like quantum anomalous Hall effect -- of the new topological materials. Among particular projects to be addressed are: 1) optical conductivity of 3D gapless Dirac fermions in the presence of smooth disorder, 2) interplay of disorder and Coulomb interactions in the spectral properties of such fermions, 3) formation and structure of the impurity band with Coulomb supercritical clusters, 4) Coulomb interaction-driven renormalization of the electron spectrum and of the transport response in the presence of strong magnetic field, 5) instanton approach to the disorder-induced fluctuation states in zero-gap 3D materials, and 6) the role of disorder in quantum anomalous Hall effect. The proposal relies upon the investigators' previous broad expertise in interacting and disordered electron systems. The methods to be employed include perturbative diagrammatic technique, non-perturbative instanton and self-consistent approximations, hydrodynamics of electron liquid. Both analytical as well as numerical approaches are to be employed. The anticipated broader outcome of the proposal includes gaining an in-depth understanding of the interplay of the disorder and interactions under the conditions when this interplay has the most dramatic impact on observables. Traditionally, interaction effects are among the most challenging and interesting problems of condensed matter physics. Similarly, disordered systems typically present very difficult but extremely rich problems in the description of various materials. Importantly, understanding the spectral and transport properties of such materials not only presents the fundamental objective, but is also of particular interest for many applications, such as computation, memory, optics, plasmonics. In particular realization of the quantum anomalous Hall effect may lead to the development of low-power-consumption electronics. Indeed, a major constraint for practical use of the quantum Hall effect is limited by the requirement of the quantizing magnetic field. At the same time, the quantum anomalous Hall effect samples exhibit non-dissipative edge quantum transport in a zero magnetic field.

36 MATERIALS SCIENCE↗

Applying Transfer Learning for Street-Scale Nuisance Flood Forecasting in Coastal-Urban Cities

An important challenge with Machine Learning (ML) is its transferability; that is, whether a ML model trained on one set of data can be applied to a second set of data without requiring a full re-training of the model. Transfer Learning (TL) addresses this challenge by transferring knowledge learned in the source domain (the data it was trained on) to the target domain (a second set of data that is statistically different but related, which the model was not trained on). This study investigates the use of TL for street-scale nuisance flood forecasting by exploring whether a ML model trained on data collected for one set of streets can effectively forecast flooding for another set of streets in the same city using TL. The envisioned use case is a city deploying a new flood depth monitoring sensor on a street and using TL to apply a ML model, trained on sensor data from an existing flood depth sensor network, to this new street. Eventually, the new flood depth sensor will have a sufficient dataset for training its own ML model, but TL can be used to fill the gap in time while this new dataset is being generated. This method is explored using a Long Short-Term Memory (LSTM) model trained on data for the flood-prone streets of Norfolk City, Virginia. The data used for training includes environmental time series (rainfall, tide), topographic features (Digital Elevation Model (DEM), Topographic Wetness Index (TWI), Depth To Water (DTW)), and street-scale flood depth time series obtained from a high-fidelity physics-based model, acting as a synthetic street-scale stream depth sensor dataset since actual stream depth sensor data is generally unavailable for most cities. A set of 180 flood-prone streets was used to train a base model, while another set of 180 flood-prone streets was used to re-train that model using different TL strategies. The results show that full-weight re-training proved most effective and minimal re-training of only the output layer was insufficient. The advantage of TL was most pronounced when target data was limited, meaning data collected at the new water depth sensor location included generally less than 18 flood events. As target data increased beyond 18 flood events, the benefit of TL diminished relative to training a ML model directly on the local flood events. These findings can assist cities as they implement street-scale flood sensing systems to create accurate forecasts for new sensing locations that do not yet have sufficient data records to train a local ML model.

Roy, Binata [Univ. of Virginia, Charlottesville, V↗

Polyakov model in ’t Hooft flux background: a quantum mechanical reduction with memory

We construct a compactification of Polyakov model on T 2 × $\mathbb{R}$ down to quantum mechanics which remembers non-perturbative aspects of field theory even at an arbitrarily small area. Standard compactification on small T 2 × $\mathbb{R}$ possesses a unique perturbative vacuum (zero magnetic flux state), separated parametrically from higher flux states, and the instanton effects do not survive in the Born-Oppenheimer approximation. By turning on a background magnetic GNO flux in co-weight lattice corresponding to a non-zero ’t Hooft flux, we show that N-degenerate vacua appear at small torus, and there are N - 1 types of flux changing instantons between them. We construct QM instantons starting with QFT instantons using the method of replicas. For example, SU(2) gauge theory with flux reduces to the double-well potential where each well is a fractional flux state. Despite the absence of a mixed anomaly, the vacuum structure of QFT and the one of QM are continuously connected. We also compare the quantum mechanical reduction of the Polyakov model with the deformed Yang-Mills, by coupling both theories to TQFTs. In particular, we compare the mass spectrum for dual photons and energy spectrum in the QM limit. We give a detailed description of critical points at infinity in the semi-classical expansion, and their role in resurgence structure.

nonperturbative effects↗

Optical and microstructural characterization of Er 3+ doped epitaxial cerium oxide on silicon

Rare-earth ion dopants in solid-state hosts are ideal candidates for quantum communication technologies, such as quantum memories, due to the intrinsic spin–photon interface of the rare-earth ion combined with the integration methods available in the solid state. Erbium-doped cerium oxide (Er:CeO 2 ) is a particularly promising host material platform for such a quantum memory, as it combines the telecom-wavelength (~ 1.5 μm) 4f–4f transition of erbium, a predicted long electron spin coherence time when embedded in CeO 2 , and a small lattice mismatch with silicon. In this work, we report on the epitaxial growth of Er:CeO 2 thin films on silicon using molecular beam epitaxy, with controlled erbium concentration between 2 and 130 parts per million (ppm). We carry out a detailed microstructural study to verify the CeO 2 host structure and characterize the spin and optical properties of the embedded Er 3+ ions as a function of doping density. In as-grown Er:CeO 2 in the 2–3 ppm regime, we identify an EPR linewidth of 245(1) MHz, an optical inhomogeneous linewidth of 9.5(2) GHz, an optical excited state lifetime of 3.5(1) ms, and a spectral diffusion-limited homogeneous linewidth as narrow as 4.8(3) MHz. We test the annealing of Er:CeO 2 films up to 900 °C, which yields narrowing of the inhomogeneous linewidth by 20% and extension of the excited state lifetime by 40%.

36 MATERIALS SCIENCE↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

Embedding Neural Thermal Scattering (NeTS) Modules in SERPENT for Higher Fidelity Advanced Reactor Analysis

When a neutron born in fission thermalizes to the order of $k$ $B$ $T$, it’s de-Broglie wavelength and energy approach the order of inter-atomic spacing and elementary lattice oscillations, respectively. $S$($a,β,t$) or the scattering law, uuantify these temperature-dependent crystallographic contributions to total cross section (or reaction rate). In a Monte Carlo analysis, cumulative distribution functions (CDFs) of $S$($a,β,t$) are loaded to memory from “A Compact ENDF” (ACE) files for stochastically selecting thermal scattered neutron trajectories. In this work, novel neural thermal scattering (NeTS) modules for $S$($a,β,t$) CDFs are designed, trained, serialized and embedded within SERPENT using Python’s limited C-API for on-the-fly deployment of crystalline graphite $S$($a,β,t$) sampling. Torchscript tracing and Numba just-in-time (JIT) compilation streamline neural inference on NVIDIA GPUs with CUDA libraries. Demonstrations of bare sphere thermalization of fast and thermal sources show excellent agreement between embedded NeTS in SERPENT and MCNP. With an explicit model of the reactor, NeTS can predict on-the-fly changes in TREAT neutron spectra as a function of local temperature, which can serve to improve transient and accident predictions in a multiphysics analysis framework. This framework can be further extended to account on-the-fly for changes in local graphitic microstructure to scattering cross sections, and outlines a novel coupling of modern machine learning with state-of-the-art reactor physics methods.

97 MATHEMATICS AND COMPUTING↗

G-Band Radar Demonstration for Microphysics Field Campaign Report

The G-Band Radar Demonstration for Microphysics (GRDM) campaign took place at the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) from March 15 to April 30, 2024. This was a deployment of two of NASA’s Jet Propulsion Laboratory (JPL) radars and one radar from Brookhaven National Laboratory to demonstrate the utility of high-frequency millimeter-wave radars for remote sensing of stratocumulus microphysical properties. The radars were deployed on the Ellen Browning Scripps Memorial Pier alongside the AMF instruments (Figure 1). The radars include a Ka-band (35 GHz), W-band (94 GHz), and four G-band (158, 165, 174, and 240 GHz) channels. The 240 GHz and W-band channels provide complete Doppler spectra, which are useful for advanced analysis. The 158-175 GHz channels are sensitive to the water vapor profile and are useful for attenuation correction. These radars complement the high-sensitivity ARM KAZR. The goal of the deployment was to observe drizzling stratocumulus and demonstrate the capabilities of the multifrequency radar data set to constrain profiles of liquid water content and drizzle drop characteristic size. The data are still being analyzed. The methodology to derive the cloud and precipitation parameters will exploit differential attenuation and differential reflectivity between low-frequency (Ka-band) and high-frequency (G-band) channels. The method will also exploit the capability of the G-band observations to constrain the attenuation due to water vapor. These observations will quantify the capabilities and limitations of the emerging technology of G-band radars for constraint stratocumulus cloud microphysics, which are key to constraining aerosol-cloud-precipitation interactions and low-cloud climate feedback.

47 OTHER INSTRUMENTATION↗

Enhancing scalability of a matrix-free eigensolver for studying many-body localization

We propose several techniques to enhance the parallel scalability of a matrix-free eigensolver designed for studying many-body localization (MBL) of quantum spin chain models with nearest-neighbor interactions and on-site disorder. This type of problem is computationally challenging because the dimension of the associated Hamiltonian matrix grows exponentially with respect to the number of spins L, and we need to average over different realizations of the random disorder to obtain relevant statistical behavior. For each disorder realization, we need to compute eigenvalues from different regions of the spectrum and their corresponding eigenvectors. In previous work, the interior eigenstates for a single eigenvalue problem are computed via the shift-and-invert Lanczos algorithm. Due to the extremely high memory footprint of the LU factorizations, this technique is not well suited for large L’s. For example, we need thousands of compute nodes on modern high performance computing infrastructures to go beyond L = 24. The matrix-free approach does not suffer from this memory bottleneck, however, its scalability is limited by a computation and communication load imbalance. To reduce this imbalance and to significantly enhance the scalability of the matrix-free eigensolver, we reorder the matrix and leverage the consistent space runtime, CSPACER. We also show its efficiency in managing irregular communication patterns at scale compared to optimized MPI non-blocking two-sided and one-sided RMA implementation variants. This effort enables us to study MBL for spin chains with a larger number of spins. The efficiency and effectiveness of the proposed algorithm is demonstrated by computing eigenstates on a massively parallel many-core high performance computer.

METIS↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer paral- lelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation.

Wang, Tianqi↗