Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory mapping”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Fusion of Geothermal and InSAR Data with Machine Learning for Enhanced Deformation Forecasting at the Geysers

The Geysers geothermal field in California is experiencing land subsidence due to the seismic and geothermal activities taking place. This poses a risk not only to the underlying infrastructure but also to the groundwater level which would reduce the water availability for the local community. Because of this, it is crucial to monitor and assess the surface deformation occurring and adjust geothermal operations accordingly. In this study, we examine the correlation between the geothermal injection and production rates as well as the seismic activity in the area, and we show the high correlation between the injection rate and the number of earthquakes. This motivates the use of this data in a machine learning model that would predict future deformation maps. First, we build a model that uses interferometric synthetic aperture radar (InSAR) images that have been processed and turned into a deformation time series using LiCSBAS, an open-source InSAR time series package, and evaluate the performance against a linear baseline model. The model includes both convolutional neural network (CNN) layers as well as long short-term memory (LSTM) layers and is able to improve upon the baseline model based on a mean squared error metric. Then, after getting preprocessed, we incorporate the geothermal data by adding them as additional inputs to the model. This new model was able to outperform both the baseline and the previous version of the model that uses only InSAR data, motivating the use of machine learning models as well as geothermal data in assessing and predicting future deformation at The Geysers as part of hazard mitigation models which would then be used as fundamental tools for informed decision making when it comes to adjusting geothermal operations.

Yazbeck, Joe (ORCID:0000000302235260)↗

Synchronization between processes in a coordination namespace

A system and method of supporting point-to-point synchronization among processes/nodes implementing different hardware barriers in a tuple space/coordinated namespace (CNS) extended memory storage architecture. The system-wide CNS provides an efficient means for storing data, communications, and coordination within applications and workflows implementing barriers in a multi-tier, multi-nodal tree hierarchy. The system provides a hardware accelerated mechanism to support barriers between the participating processes. Also architected is a tree structure for a barrier processing method where processes are mapped to nodes of a tree, e.g., a tree of degree k to provide an efficient way of scaling the number of processes in a tuple space/coordination namespace.

Jacob, Philip↗

Synchronization between processes in a coordination namespace

A system and method of supporting point-to-point synchronization among processes/nodes implementing different hardware barriers in a tuple space/coordinated namespace (CNS) extended memory storage architecture. The system-wide CNS provides an efficient means for storing data, communications, and coordination within applications and workflows implementing barriers in a multi-tier, multi-nodal tree hierarchy. The system provides a hardware accelerated mechanism to support barriers between the participating processes. Also architected is a tree structure for a barrier processing method where processes are mapped to nodes of a tree, e.g., a tree of degree k, to provide an efficient way of scaling the number of processes in a tuple space/coordination namespace.

Jacob, Philip↗

Fourier-based three-dimensional multistage transformer for aberration correction in multicellular specimens

High-resolution tissue imaging is often compromised by sample-induced optical aberrations that degrade resolution and contrast. Although wavefront sensor-based adaptive optics (AO) can measure these aberrations, such hardware solutions are typically complex, expensive to implement and slow when serially mapping spatially varying aberrations across large fields of view. Here we introduce AOViFT (adaptive optical vision Fourier transformer)—a machine learning-based aberration sensing framework built around a three-dimensional multistage vision transformer that operates on Fourier domain embeddings. AOViFT infers aberrations and restores diffraction-limited performance in puncta-labeled specimens with substantially reduced computational cost, training time and memory footprint compared to conventional architectures or real-space networks. We validated AOViFT on live gene-edited zebrafish embryos, demonstrating its ability to correct spatially varying aberrations using either a deformable mirror or postacquisition deconvolution. By eliminating the need for the guide star and wavefront sensing hardware and simplifying the experimental workflow, AOViFT lowers technical barriers for high-resolution volumetric microscopy across diverse biological samples.

Alshaabi, Thayer [Howard Hughes Medical Institute,↗

Black hole explosions as probes of new physics

The final stage of black hole evaporation is a potent probe of physics beyond the Standard Model: Hawking-Bekenstein radiation may be affected by quantum gravity “memory burden effects,” or by the presence of “dark,” beyond-the-Standard-Model degrees of freedom in ways that are testable with high-energy gamma-ray observations. We argue that information on either scenario can best be inferred from measurements of the evaporation’s light curve and by correlating observations at complementary energies. We offer several new analytical insights in how such observations map on the fundamental properties of the evaporating black holes and of the possible exotic particles they can evaporate into. Published by the American Physical Society 2025

Federico, Kevin (ORCID:0009000708441977)↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Computing the Properties of Matter with Leadership Computing Resources (Closeout Report for DE-SC0018121)

In order to add more capabilities to Halide, we have designed a new framework called Tiramisu and integrated this framework into Halide. Since Tiramisu enables Halide to target heterogeneous architectures, our development efforts have been refocused on Tiramisu. Most high-performance computer systems today are complex and increasingly heterogeneous; they may have CPUs, GPUs and FPGAs. Achieving best performance requires taking full advantage of all these different architectures. To address this issue, we have designed Tiramisu, an optimization framework that enables Halide (and other DSLs) to target heterogeneous architectures. Tiramisu is an optimization framework that takes as input a high level, architecture-independent representation of code and a set of scheduling and data mapping commands that guide code transformation. The input can either be generated by a domain-specific language (DSL) compiler such as Halide or directly written by a programmer. Tiramisu then applies the user-specified code and data-layout transformations and generates an architecture-specific, low-level intermediate representation (IR) that takes advantage of modern architectural features such as multicore parallelism, non-uniform memory (NUMA) hierarchies, clusters, and accelerators like GPUs and FPGAs. We integrated Tiramisu within Halide and implemented a representative set of benchmarks to evaluate this integration. Tiramisu is now open source and is available for public use (http://tiramisu-compiler.org/). A paper about Tiramisu was published, it shows that Tiramisu extends Halide with many new capabilities and that Tiramisu can generate efficient code for multicores, GPUs, FPGAs and distributed heterogeneous systems. The performance of code generated by the Tiramisu backends matches or exceeds hand optimized reference implementations. For example, the multicore backend matches the highly optimized Intel MKL library on many kernels and shows speedups reaching 4x over the original Halide. In addition to making Tiramisu more robust, we have used Tiramisu to implement a set of representative tensor operation for constructing baryon building blocks required for multi baryon contractions in LQCD. In order to implement this code, we needed to generalize Tiramisu in two ways: first we needed to support indirect array accesses, and second, we needed to add support for complex numbers to Tiramisu. The code generated by Tiramisu is 6x faster than the reference code. Our efforts towards an MPI based multi-node version of tiramisu have matured and the resulting code scales well on multiple nodes (tests up to 512 KNL nodes have been undertaken).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Synoptic VLBI Technique for Localizing Nonrepeating Fast Radio Bursts with CHIME/FRB

We demonstrate the blind interferometric detection and localization of two fast radio bursts (FRBs) with subarcminute precision on the 400 m baseline between the Canadian Hydrogen Intensity Mapping Experiment (CHIME) and the CHIME Pathfinder. In the same spirit as Very Long Baseline Interferometry (VLBI), the telescopes were synchronized to separate clocks, and the channelized voltage (herein referred to as baseband) data were saved to a disk with correlation performed offline. The simultaneous wide field of view and high sensitivity required for blind FRB searches implies a high data rate—6.5 terabits per second (Tb/s) for CHIME and 0.8 Tb s{sup −1} for the Pathfinder. Since such high data rates cannot be continuously saved, we buffer data from both telescopes locally in memory for ≈40 s, and write to the disk upon receipt of a low-latency trigger from the CHIME Fast Radio Burst Instrument (CHIME/FRB). The ≈200 deg{sup 2} field of view of the two telescopes allows us to use in-field calibrators to synchronize the two telescopes without needing either separate calibrator observations or an atomic timing standard. In addition to our FRB observations, we analyze bright single pulses from the pulsars B0329+54 and B0355+54 to characterize systematic localization errors. Our results demonstrate the successful implementation of key software, triggering, and calibration challenges for CHIME/FRB Outriggers: cylindrical VLBI outrigger telescopes which, along with the CHIME telescope, will localize thousands of single FRB events with sufficient precision to unambiguously associate a host galaxy with each burst.

47 OTHER INSTRUMENTATION↗

Colossal Strain Tuning of Ferroelectric Transitions in KNbO 3 Thin Films

Strong coupling between polarization ( P ) and strain (ɛ) in ferroelectric complex oxides offers unique opportunities to dramatically tune their properties. Here colossal strain tuning of ferroelectricity in epitaxial KNbO 3 thin films grown by sub‐oxide molecular beam epitaxy is demonstrated. While bulk KNbO 3 exhibits three ferroelectric transitions and a Curie temperature ( T c ) of ≈676 K, phase‐field modeling predicts that a biaxial strain of as little as −0.6% pushes its T c > 975 K, its decomposition temperature in air, and for −1.4% strain, to T c > 1325 K, its melting point. Furthermore, a strain of −1.5% can stabilize a single phase throughout the entire temperature range of its stability. A combination of temperature‐dependent second harmonic generation measurements, synchrotron‐based X‐ray reciprocal space mapping, ferroelectric measurements, and transmission electron microscopy reveal a single tetragonal phase from 10 K to 975 K, an enhancement of ≈46% in the tetragonal phase remanent polarization ( P r ), and a ≈200% enhancement in its optical second harmonic generation coefficients over bulk values. These properties in a lead‐free system, but with properties comparable or superior to lead‐based systems, make it an attractive candidate for applications ranging from high‐temperature ferroelectric memory to cryogenic temperature quantum computing.

36 MATERIALS SCIENCE↗

Short-Term Rainfall Prediction Based on Radar Echo Using an Improved Self-Attention PredRNN Deep Learning Model

Accurate short-term precipitation forecast is extremely important for urban flood warning and natural disaster prevention. In this paper, we present an innovative deep learning model named ISA-PredRNN (improved self-attention PredRNN) for precipitation nowcasting based on radar echoes on the basis of the advanced PredRNN-V2. We introduce the self-attention mechanism and the long-term memory state into the model and design a new set of gating mechanisms. To better capture different intensities of precipitation, the loss function with weights was designed. We further train the model using a combination of reverse scheduled sampling and scheduled sampling to learn the long-term dynamics from the radar echo sequences. Experimental results show that the new model (ISA-PredRNN) can effectively extract the spatiotemporal features of radar echo maps and obtain radar echo prediction results with a small gap from the ground truths. From the comparison with the other six models, the new ISA-PredRNN model has the most accurate prediction results with a critical success index (CSI) of 0.7001, 0.5812 and 0.3052 under the radar echo thresholds of 10 dBZ, 20 dBZ and 30 dBZ, respectively.

Wu, Dali (ORCID:0000000231177074)↗

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng↗

Conformational spread drives the evolution of the calcium–calmodulin protein kinase II

The calcium calmodulin (Ca 2+ /CaM) dependent protein kinase II (CaMKII) decodes Ca 2+ frequency oscillations. The CaMKIIα isoform is predominantly expressed in the brain and has a central role in learning. I matched residue and organismal evolution with collective motions deduced from the atomic structure of the human CaMKIIα holoenzyme to learn how its ring architecture abets function. Protein dynamic simulations showed its peripheral kinase domains (KDs) are conformationally coupled via lateral spread along the central hub. The underlying β-sheet motions in the hub or association domain (AD) were deconvolved into dynamic couplings based on mutual information. They mapped onto a coevolved residue network to partition the AD into two distinct sectors. A second, energetically stressed sector was added to ancient bacterial enzyme dimers for assembly of the ringed hub. The continued evolution of the holoenzyme after AD–KD fusion targeted the sector’s ring contacts coupled to the KD. Among isoforms, the α isoform emerged last and, it alone, mutated rapidly after the poikilotherm–homeotherm jump to match the evolution of memory. The correlation between dynamics and evolution of the CaMKII AD argues single residue substitutions fine-tune hub conformational spread. The fine-tuning could increase CaMKIIα Ca 2+ frequency response range for complex learning functions.

59 BASIC BIOLOGICAL SCIENCES↗

Thermal Radiation Transport with Tensor Trains

We present a novel tensor network algorithm to solve the time-dependent, gray thermal radiation transport equation. The method invokes a tensor train (TT) decomposition for the specific intensity. The efficiency of this approach is dictated by the rank of the decomposition. When the solution is “low rank,” the memory footprint of the specific intensity solution vector may be significantly compressed. The algorithm, following a step-then-truncate approach of a traditional discrete ordinates method, operates directly on the compressed state vector, thereby enabling large speedups for low-rank solutions. To achieve these speedups, we rely on a recently developed rounding approach based on the Gram-SVD. We detail how familiar S N algorithms for (gray) thermal transport can be mapped to this TT framework and present several numerical examples testing both the optically thick and thin regimes. The TT framework finds low-rank structure and supplies up to ≃60× speedups and ≃1000× compressions for problems demanding large angle counts, thereby enabling previously intractable SN calculations and supplying a promising avenue to mitigate ray effects.

79 ASTRONOMY AND ASTROPHYSICS↗

Decoherence by warm horizons

Recently Danielson, Satishchandran, and Wald (DSW) have shown that quantum superpositions held outside of Killing horizons will decohere at a steady rate. This occurs because of the inevitable radiation of soft photons (gravitons), which imprint a electromagnetic (gravitational) “which-path” memory onto the horizon. Rather than appealing to this global description, an experimenter ought to also have a local description for the cause of decoherence. One might intuitively guess that this is just the bombardment of Hawking/Unruh radiation on the system, however simple calculations challenge this idea—the same superposition held in a finite temperature inertial laboratory does not decohere at the DSW rate. In this work we provide a local description of the decoherence by mapping the DSW setup onto a worldline-localized model resembling an Unruh-DeWitt particle detector. We present an interpretation in terms of random local forces which do not sufficiently self-average over long times. Using the Rindler horizon as a concrete example we clarify the crucial role of temperature, and show that the Unruh effect is the only quantum mechanical effect underlying these random forces. A general lesson is that for an environment which induces Ohmic friction on the central system (as one gets from the classical Abraham-Lorentz-Dirac force, in an accelerating frame) the fluctuation-dissipation theorem implies that when this environment is at finite temperature it will cause steady decoherence on the central system. Our results agree with DSW and provide the complementary local perspective. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Stimuli-responsive petroleum cement composite with giant expansion and enhanced mechanical properties

Shrinkage, as an inherent property of cement during the curing process, has been a concern in the oil & gas industry for centuries. Therefore, expansive cement has been studied for many decades, primarily through incorporation of expansive additives. Nevertheless, traditional expansive additives are not able to achieve the required expansive ratio or maintain a good mechanical property for the cement composite at extreme underground condition. Accordingly, a new generation of expansive additive, which not only has the required expansion, but also work at the downhole high temperature and high pressure environment, and maintain or even enhance the mechanical property, is highly desired. In this work, a high enthalpy storage shape memory polymer (SMP) as a new class of expansive additive is investigated. It is found that the cement composite achieves 1.4% circumferential expansion by only 6% by weight of SMP additives. The compressive strength and flexural strength are also enhanced at the same time, which is hardly achievable by other expansive additives with the same concertation. Good pumpability is also proved by rheological study. The mechanism controlling the enhanced properties is reveled through morphological study and element mapping analysis at the SMP/cement interface.

36 MATERIALS SCIENCE↗

Interfacial-Strain-Controlled Ferroelectricity in Self-Assembled BiFeO 3 Nanostructures

Self-assembled BiFeO 3 -CoFe 2 O 4 (BFO-CFO) vertically aligned nanocomposites are promising for logic, memory, and multiferroic applications, primarily due to the tunability enabled by strain engineering at the prodigious epitaxial vertical interfaces. However, local investigations directly revealing functional properties in the vicinity of such critical interfaces are often hampered by the size, geometry, microstructure, and concomitant experimental artifacts. Ferroelectric switching in the presence of lateral distributions of vertical strain thus remains relatively unexplored, with broader implications for all strain-engineered functional devices. Additionally, by implementing tomographic atomic force microscopy, 3D domain orientation mapping, and spatially-resolved ferroelectric switching movies, local tensile strain significantly impacts the ferroelectric switching, principally by retarding domain nucleation in the BFO nearest to the vertically epitaxial tensile-strained interfaces. The relaxed centers of the BFO pillars become preferred domain nucleation and growth sites for low biases, with up to an order of magnitude change in the edge:center switching ratio for high biases. The new, multi-dimensional imaging approach—and its corresponding insights especially for directly strained interface effects on local properties—thereby advances the fundamental understanding of polarization switching and provides design principles for optimizing functional response in confined nanoferroic systems.

36 MATERIALS SCIENCE↗

Pre-conditioned BFGS-based uncertainty quantification in elastic full-waveform inversion

SUMMARY Full-waveform inversion has become an essential technique for mapping geophysical subsurface structures. However, proper uncertainty quantification is often lacking in current applications. In theory, uncertainty quantification is related to the inverse Hessian (or the posterior covariance matrix). Even for common geophysical inverse problems its calculation is beyond the computational and storage capacities of the largest high-performance computing systems. In this study, we amend the Broyden–Fletcher–Goldfarb–Shanno (BFGS) algorithm to perform uncertainty quantification for large-scale applications. For seismic inverse problems, the limited-memory BFGS (L-BFGS) method prevails as the most efficient quasi-Newton method. We aim to augment it further to obtain an approximate inverse Hessian for uncertainty quantification in FWI. To facilitate retrieval of the inverse Hessian, we combine BFGS (essentially a full-history L-BFGS) with randomized singular value decomposition to determine a low-rank approximation of the inverse Hessian. Setting the rank number equal to the number of iterations makes this solution efficient and memory-affordable even for large-scale problems. Furthermore, based on the Gauss–Newton method, we formulate different initial, diagonal Hessian matrices as pre-conditioners for the inverse scheme and compare their performances in elastic FWI applications. We highlight our approach with the elastic Marmousi benchmark model, demonstrating the applicability of pre-conditioned BFGS for large-scale FWI and uncertainty quantification.

58 GEOSCIENCES↗

Nanoscale Probing of Electrical Memory Effects in van der Waals Layered PdSe 2

Tunable electronic materials that can be switched between different impedance states are fundamental to the hardware elements for neuromorphic computing architectures. This “brain-like” computing paradigm uses highly paralleled and colocated data processing, leading to greatly improved energy efficiency and performance compared to traditional architectures in which data have to be frequently transferred between processor and memory. In this work, we use scanning microwave impedance microscopy for nanoscale electrical and electronic characterization of two-dimensional layered semiconductor PdSe 2 to probe neuromorphic properties. The local resolution of tens of nanometers reveals significant differences in electronic behavior between and within PdSe 2 nanosheets (NSs). In particular, we detected both n-type and p-type behaviors, although previous reports only point to ambipolar n-type dominating characteristics. Nanoscale capacitance–voltage curves and subsequent calculation of characteristic maps revealed a hysteretic behavior originating from the creation and erasure of Se vacancies as well as the switching of defect charge states. In addition, stacks consisting of two NSs show enhanced resistive and capacitive switching, which is attributed to trapped charge carriers at the interfaces between the stacked NSs. Stacking n- and p-type NSs results in a combined behavior that allows one to tune electrical characteristics. In conclusion, as local inhomogeneities of electrical and electronic behavior can have a significant impact on the overall device performance, the demonstrated nanoscale characterization and analysis will be applicable to a wide range of semiconducting materials.

2D material↗