Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory cell”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Threadsafe Dynamic Neighbor Lists for Monte Carlo Ray Tracing

Monte Carlo (MC) transport codes offer high-fidelity modeling of particle transport physics, but their high computational cost makes them impractical for many applications. For some applications such as multiphysics and depletion that use finely discretized geometries, a large portion of this computational cost is attributable to ray tracing. Neighbor lists are a well-known method for accelerating ray-tracing calculations in a MC code, but despite their prevalence, little work has been published on the details of their implementation. The fine details can have a significant impact on performance, particularly when using shared-memory parallelism. This paper addresses these details of implementation with a discussion of different neighbor list schemes and their impact on software runtime. Performance tests were run by using OpenMC on a pin-cell problem discretized with up to 200 axial regions. The results demonstrate that switching from surface-based to cell-based neighbor lists leads to a 10 faster calculation rate for the most fine discretization. Finally, using a threadsafe shared-memory data structure results in a 20% faster calculation rate versus simple threadprivate neighbor lists. Results here show that a data structure that is contiguous in memory improves performance by only 1% to 2% over noncontiguous linked lists.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

EAP Patterns

EAP Patterns: Memory access and iteration patterns from the EAP code base with the physics removed. This is intended to be a serial app representing memory access patterns. We will populate the data structures from EAP output files to provide representative patterns of face and cell loops within the EAP code base The arguments to the program are the EAP output file and the number of MPI processors to emulate. Currently there is no MPI in the application and the number of processors represents a means of emulating the halo (clone) cells around the domain of the given processor. An optional argument `processor_ID` can be provided to specify the processor to emulate. In the future we plan to include the ability to run multiple MPI ranks to more accurately represent the on-node demands on memory bandwidth.

Swaminarayan, Sriram↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗

The human aortic endothelium undergoes dose-dependent DNA methylation in response to transient hyperglycemia

Glycemic control is a strong predictor of long-term cardiovascular risk in patients with diabetes mellitus, and poor glycemic control influences long-term risk of cardiovascular disease even decades after optimal medical management. This phenomenon, termed glycemic memory, has been proposed to occur due to stable programs of cardiac and endothelial cell gene expression. This transcriptional remodeling has been shown to occur in the vascular endothelium through a yet undefined mechanism of cellular reprogramming.

60 APPLIED LIFE SCIENCES↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

Harnessing ferro-valleytricity in pentalayer rhombohedral graphene for memory and compute

Two-dimensional materials with multiple degrees of freedom, including spin, valleys, and orbitals, open up an exciting avenue for engineering multifunctional devices. Beyond spintronics, these degrees of freedom can lead to novel quantum effects such as valley-dependent Hall effects and orbital magnetism, which could revolutionize next-generation electronics. However, achieving independent control over valley polarization and orbital magnetism has been a challenge due to the need for large electric fields. A recent breakthrough involving pentalayer rhombohedral graphene has demonstrated the ability to individually manipulate anomalous Hall signals and orbital magnetic hysteresis, forming what is known as a valley-magnetic quartet. Here, we leverage the electrically tunable ferro-valleytricity of pentalayer rhombohedral graphene to develop nonvolatile memory and in-memory computation applications. We propose an architecture for a dense, scalable, and selector-less nonvolatile memory array that harnesses the electrically tunable ferro-valleytricity. In our designed array architecture, nondestructive read and write operations are conducted by sensing the valley state through two different pairs of terminals, allowing for independent optimization of read/write peripheral circuits. The power consumption of our PRG-based array is remarkably low, with only ∼6 nW required per write operation and ∼2.3 nW per read operation per cell. This consumption is orders of magnitude lower than that of the majority of state-of-the-art cryogenic memories. Additionally, we engineer in-memory computation by implementing majority logic operations within our proposed nonvolatile memory array without modifying the peripheral circuitry. In conclusion, our framework presents a promising pathway toward achieving ultra-dense cryogenic memory and in-memory computation capabilities.

2D materials↗

Association of AK4 Protein From Stem Cell–Derived Neurons With Cognitive Reserve

Background and Objectives: Identifying protein targets that provide cognitive reserve is a strategy to prevent and treat Alzheimer disease and Alzheimer disease related dementias (AD/ADRD). Previous studies using bulk human brain tissue reported 12 proteins associated with cognitive reserve. This study examined whether the same proteins from induced neurons (iNs) are associated with cognitive reserve of their human donors. Methods: Here, induced pluripotent stem cell (iPSC) lines were generated from cryopreserved peripheral blood mononuclear cells of older adults who were autopsied as part of the Religious Orders Study or Rush Memory and Aging Project. Neurons were induced from iPSCs using a standard neurogenin2 protocol. Tandem mass tag proteomics analyses were conducted on iNs day 21. Cognitive reserve of their human donors was measured as person-specific slopes of cognitive change not accounted for by common neuropathologies. Results: The 53 human donors died at a mean age of 91 years, all were non-Latino White, and 36 (67.9%) were female. Eighteen were diagnosed with Alzheimer dementia proximate to death, and 34 had pathologic AD diagnosis at autopsy. Approximately 60% of the donors had above-average cognitive reserve such that their cognition declined slower than an average person with comparable burdens of neuropathologies. Eight of the 12 candidate proteins were quantified in iNs proteomics analyses. Higher adenylate kinase 4 (AK4) expression in iNs was associated with lower cognitive reserve, consistent with the previous report for brain AK4 expression. Discussion: By replicating cortical protein associations with cognitive reserve in human iNs, these data provide a valuable molecular readout for studying complex clinical phenotypes such as cognitive reserve in a dish.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing Lattice Kinetic Schemes for Fluid Dynamics with Lattice-Equivariant Neural Networks

A new class of equivariant neural networks is presented, hereby dubbed lattice-equivariant neural networks (LENNs), designed to satisfy local symmetries of a lattice structure. The approach develops within a recently introduced framework aimed at learning neural network-based surrogate models’ lattice Boltzmann collision operators. Whenever neural networks are employed to model physical systems, respecting symmetries and equivariance properties has been shown to be key for accuracy, numerical stability, and performance. Here, hinging on ideas from group representation theory, trainable layers are defined whose algebraic structure is equivariant with respect to the symmetries of the lattice cell. In this work, the presented method naturally allows for efficient implementations, in terms of both memory usage and computational costs, supporting scalable training/testing for lattices in two spatial dimensions and higher (in which the size of symmetry group grows). The approach is validated and tested considering 2D and 3D flowing dynamics, both in laminar and turbulent regimes. It is compared with group-averaged-based symmetric networks and with plain, nonsymmetric, networks, showing how the presented approach unlocks the (a posteriori) accuracy and training stability of the former models and the train/inference speed of the latter networks. (LENNs are about one order of magnitude faster than group-averaged networks in 3D.) The work in this paper opens toward practical use of machine learning-augmented lattice Boltzmann CFD in real-world simulations.

97 MATHEMATICS AND COMPUTING↗

A PIPS + SrI 2 (Eu) detector for atmospheric radioxenon monitoring

The PIPS–SrI 2 (Eu) is a prototype atmospheric radioxenon detection system designed at Oregon State University in support of international efforts towards monitoring clandestine nuclear weapon testing activities. This detector aims to address some shortcomings found in currently deployed beta–gamma atmospheric radioxenon detection systems, such as lackluster energy resolution and memory effect, by employing modern detection materials and readout. The system uses a PIPSBox, a silicon-based gas cell, for electron detection, and a pair of ultrabright, D-shaped SrI 2 (Eu) scintillators coupled to silicon photomultipliers for photon detection. A custom eight-channel digital pulse processor equipped with a field programmable gate-array (FPGA) identifies electron–photon coincidences between the volumes in near real-time. Gas samples of the four radioxenon isotopes of interest were independently measured with the PIPS–SrI 2 (Eu) detection system to determine energy resolution and efficiency. Application of FPGA-based coincidence discrimination in near real-time reduced the ambient background count rate by 95.85 ± 0.04%. Using parameters from the Xenon International gas processing unit and assuming a blank sample and zero memory effect the minimum detectable concentrations (MDCs) for the isotopes were calculated to be 0.12 ± 0.03, 0.27 ± 0.05, 0.15 ± 0.02, and 1.00 ± 0.08 mBq/m 3 air for 131m Xe, 133 Xe, 133m Xe, and 135 Xe, respectively. These MDC estimates compare well with other radioxenon detection systems employed in the International Monitoring System (IMS) and indicate that the PIPS–SrI 2 (Eu) is in compliance with the Comprehensive Nuclear Test-Ban-Treaty Organization (CTBTO) sensitivity requirement of ≤ 1 mBq/m 3 for 133 Xe.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory

The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.

59 BASIC BIOLOGICAL SCIENCES↗

A review of materials used in tomographic volumetric additive manufacturing

Abstract Volumetric additive manufacturing is a novel fabrication method allowing rapid, freeform, layer-less 3D printing. Analogous to computer tomography (CT), the method projects dynamic light patterns into a rotating vat of photosensitive resin. These light patterns build up a three-dimensional energy dose within the photosensitive resin, solidifying the volume of the desired object within seconds. Departing from established sequential fabrication methods like stereolithography or digital light printing, volumetric additive manufacturing offers new opportunities for the materials that can be used for printing. These include viscous acrylates and elastomers, epoxies (and orthogonal epoxy-acrylate formulations with spatially controlled stiffness) formulations, tunable stiffness thiol-enes and shape memory foams, polymer derived ceramics, silica-nanocomposite based glass, and gelatin-based hydrogels for cell-laden biofabrication. Here we review these materials, highlight the challenges to adapt them to volumetric additive manufacturing, and discuss the perspectives they present. Graphical abstract

36 MATERIALS SCIENCE↗

Properties of Electronic Materials

This final technical report summarizes the research conducted under DOE Grant DE-SC0002623, "Properties of Electronic Materials," led by Principal Investigator Shengbai Zhang at Rensselaer Polytechnic Institute. Over the 16-year period, the project employed first-principles computational methods to investigate the structural, electronic, and dynamic properties of a wide range of electronic materials, with applications in energy technologies, optoelectronics, and data storage. Key areas included topological insulators, phase-change materials, graphene and two-dimensional systems, perovskites for photovoltaics, defect engineering in semiconductors, kagome lattices, and ultrafast carrier dynamics. The research resulted in 115 peer-reviewed publications, advancing fundamental understanding of material behaviors at the atomic scale and contributing to innovations in renewable energy, memory devices, and quantum materials. Findings have implications for improving energy efficiency, developing lead-free solar cells, and enabling high-speed data processing. The work has trained numerous graduate students and postdocs, fostering the next generation of computational materials scientists. The original goals were to develop theoretical models and computational tools to predict and optimize electronic properties of materials for energy applications. All objectives were accomplished, with no major departures from planned methodologies. Challenges in computational scaling were addressed through access to high-performance computing resources.

36 MATERIALS SCIENCE↗

Loss of Mitochondrial Tusc2/Fus1 Triggers a Brain Pro-Inflammatory Microenvironment and Early Spatial Memory Impairment

Brain pathological changes impair cognition early in disease etiology. There is an urgent need to understand aging-linked mechanisms of early memory loss to develop therapeutic strategies and prevent the development of cognitive impairment. Tusc2 is a mitochondrial-resident protein regulating Ca 2+ fluxes to and from mitochondria impacting overall health. We previously reported that Tusc2 –/– female mice develop chronic inflammation and age prematurely, causing age- and sex-dependent spatial memory deficits at 5 months old. Therefore, we investigated Tusc2-dependent mechanisms of memory impairment in 4-month-old mice, comparing changes in resident and brain-infiltrating immune cells. Interestingly, Tusc2 –/– female mice demonstrated a pro-inflammatory increase in astrocytes, expression of IFN-γ in CD4 + T cells and Granzyme-B in CD8 + T cells. We also found fewer FOXP3 + T-regulatory cells and Ly49G + NK and Ly49G + NKT cells in female Tusc2 –/– brains, suggesting a dampened anti-inflammatory response. Moreover, Tusc2 –/– hippocampi exhibited Tusc2- and sex-specific protein changes associated with brain plasticity, including mTOR activation, and Calbindin and CamKII dysregulation affecting intracellular Ca2 + dynamics. Overall, the data suggest that dysregulation of Ca 2+ -dependent processes and a heightened pro-inflammatory brain microenvironment in Tusc2 –/– mice could underlie cognitive impairment. Thus, strategies to modulate the mitochondrial Tusc2- and Ca 2+ - signaling pathways in the brain should be explored to improve cognitive health.

59 BASIC BIOLOGICAL SCIENCES↗

Materials and methods for the preparation of nanocomposites

Disclosed herein is an isolable colloidal particle comprising a nanoparticle and an inorganic capping agent bound to the surface of the nanoparticle, a method for making the same in a biphasic solvent mixture, and the formation of structures and solids from the isolable colloidal particle. The process can yield photovoltaic cells, piezoelectric crystals, thermoelectric layers, optoelectronic layers, light emitting diodes, ferroelectric layers, thin film transistors, floating gate memory devices, phase change layers, and sensor devices.

36 MATERIALS SCIENCE↗

On the Feasibility of 1T Ferroelectric FET Memory Array

To fully exploit the ferroelectric field effect transistor (FeFET) as compact embedded nonvolatile memory for various computing and storage applications, it is desirable to use a single FeFET (1T) as a unit cell and arrange the cells into an array. However, many write mechanisms for an 1T FeFET array reported in the literature are yet to be validated experimentally. In this work, we performed a comprehensive experimental characterization on the write operations in an 1T- NOR and 1T- AND array using n-channel bulk FeFETs. We discovered that: 1) the source/drain contact can only supply minority carriers (i.e., electrons) to the channel for polarization screening during the low- V TH state programming; 2) the body contact can not only supply majority carriers (i.e., holes) for efficient high- V TH state write, but also depletion charge for low- V TH state programming, though with lower efficiency; and 3) during the low/high- V TH programming, only the path that can supply negative/positive screening charges, respectively, need to respond, which necessitates the application of proper write biases on the corresponding terminal. Based on the understanding of these write mechanisms, we show the importance of localized body contact or column-wise body contact for the successful high- V TH state programming. We also show that the previously proposed C- AND write scheme fails to program the FeFET to the low- V TH state for our devices. Finally, we propose several write schemes for both 1T- AND and 1T- NOR arrays for various scenarios providing insights for choosing the appropriate write scheme, which will facilitate the adoption of 1T FeFET memory arrays for emerging applications.

FeFET↗

HiPACE++: A portable, 3D quasi-static particle-in-cell code

Modeling plasma accelerators is a computationally challenging task and the quasi-static particle-in-cell algorithm is a method of choice in a wide range of situations. In this work, we present the first performance-portable, quasi-static, three-dimensional particle-in-cell code HiPACE++. By decomposing all the computation of a 3D domain in successive 2D transverse operations and choosing appropriate memory management, HiPACE++ demonstrates orders-of-magnitude speedups on modern scientific GPUs over CPU-only implementations. The 2D transverse operations are performed on a single GPU, avoiding time-consuming communications. The longitudinal parallelization is done through temporal domain decomposition, enabling near-optimal strong scaling from 1 to 512 GPUs. HiPACE++ is a modular, open-source code enabling efficient modeling of plasma accelerators from laptops to state-of-the-art supercomputers.

97 MATHEMATICS AND COMPUTING↗