Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sylvester-preconditioned adaptive-rank implicit time integrators for advection-diffusion equations with variable coefficients

Here, we consider the adaptive-rank integration of multi-dimensional time-dependent advection-diffusion partial differential equations (PDEs) with variable coefficients. We employ a standard finite-difference method for spatial discretization coupled with high-order diagonally implicit Runge-Kutta temporal schemes. The discrete equation is a generalized Sylvester equation (GSE), which we solve with a projection-based adaptive-rank algorithm structured around two key strategies: (i) constructing dimension-wise subspaces using a novel atypical extended Krylov strategy, and (ii) efficiently solving the basis coefficient matrix with a preconditioned GMRES solver. The low-rank decomposition is performed in 2D using SVD and with high-order SVD (HOSVD) in 3D to represent the tensor in a compressed Tucker format. For d-dimensional problems (here, d = 2 or 3), the computational complexity and memory storage of the approach are found numerically to scale as and $\mathscr{O}(Nr^2) + \mathscr{O} (r^{d+1})$ and $\mathscr{O}(Nr) + \mathscr{O} (r^{d})$, respectively, with the one-dimensional resolution and the maximal rank during the Krylov iteration (which we find to be largely independent of on our numerical examples). We present numerical examples that illustrate the advertised properties of the algorithm.

97 MATHEMATICS AND COMPUTING↗

Electroencephalographic imaging of higher brain function

High temporal resolution is necessary to resolve the rapidly changing patterns of brain activity that underlie mental function. Electroencephalography (EEG) provides temporal resolution in the millisecond range. However, traditional EEG technology and practice provide insufficient spatial detail to identify relationships between brain electrical events and structures and functions visualized by magnetic resonance imaging or positron emission tomography. Recent advances help to overcome this problem by recording EEGs from more electrodes, by registering EEG data with anatomical images, and by correcting the distortion caused by volume conduction of EEG signals through the skull and scalp. In addition, statistical measurements of sub-second interdependences between EEG time-series recorded from different locations can help to generate hypotheses about the instantaneous functional networks that form between different cortical regions during perception, thought and action. Example applications are presented from studies of language, attention and working memory. Along with its unique ability to monitor brain function as people perform everyday activities in the real world, these advances make modern EEG an invaluable complement to other functional neuroimaging modalities.

Non-NASA Center↗

Process and representation in graphical displays

How people comprehend graphics is examined. Graphical comprehension involves the cognitive representation of information from a graphic display and the processing strategies that people apply to answer questions about graphics. Research on representation has examined both the features present in a graphic display and the cognitive representation of the graphic. The key features include the physical components of a graph, the relation between the figure and its axes, and the information in the graph. Tests of people's memory for graphs indicate that both the physical and informational aspect of a graph are important in the cognitive representation of a graph. However, the physical (or perceptual) features overshadow the information to a large degree. Processing strategies also involve a perception-information distinction. In order to answer simple questions (e.g., determining the value of a variable, comparing several variables, and determining the mean of a set of variables), people switch between two information processing strategies: (1) an arithmetic, look-up strategy in which they use a graph much like a table, looking up values and performing arithmetic calculations; and (2) a perceptual strategy in which they use the spatial characteristics of the graph to make comparisons and estimations. The user's choice of strategies depends on the task and the characteristics of the graph. A theory of graphic comprehension is presented.

Gillan, Douglas J.↗

Exocomets from a Solar System Perspective

Exocomets are small bodies releasing gas and dust which orbit stars other than the Sun. Their existence was first inferred from the detection of variable absorption features in stellar spectra in the late 1980s using spectroscopy. More recently, they have been detected through photometric transits from space, and through far-IR/mm gas emission within debris disks. As (exo)comets are considered to contain the most pristine material accessible in stellar systems, they hold the potential to give us information about early stage formation and evolution conditions of extra solar systems. In the solar system, comets carry the physical and chemical memory of the protoplanetary disk environment where they formed, providing relevant information on processes in the primordial solar nebula. The aim of this paper is to compare essential compositional properties between solar system comets and exocomets to allow for the development of new observational methods and techniques. The paper aims to highlight commonalities and to discuss differences which may aid the communication between the involved research communities and perhaps also avoid misconceptions. The compositional properties of solar system comets and exocomets are summarized before providing an observational comparison between them. Exocomets likely vary in their composition depending on their formation environment like solar system comets do, and since exocomets are not resolved spatially, they pose a challenge when comparing them to high fidelity observations.

comets↗

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

Particle hit clustering and identification using point set transformers in liquid argon time projection chambers

Liquid argon time projection chambers are often used in neutrino physics and dark-matter searches because of their high spatial resolution. The images generated by these detectors are extremely sparse, as the energy values detected by most of the detector are equal to 0, meaning that despite their high resolution, most of the detector is unused in a particular interaction. Instead of representing all of the empty detections, the interaction is usually stored as a sparse matrix, a list of detection locations paired with their energy values. Traditional machine learning methods that have been applied to particle reconstruction such as convolutional neural networks (CNNs), however, cannot operate over data stored in this way and therefore must have the matrix fully instantiated as a dense matrix. Operating on dense matrices requires a lot of memory and computation time, in contrast to directly operating on the sparse matrix. We propose a machine learning model using a point set neural network that operates over a sparse matrix, greatly improving both processing speed and accuracy over methods that instantiate the dense matrix, as well as over other methods that operate over sparse matrices. Compared to competing state-of-the-art methods, our method improves classification performance by 14%, segmentation performance by more than 22%, while taking 80% less time and using 66% less memory. Compared to state-of-the-art CNN methods, our method improves classification performance by more than 86%, segmentation performance by more than 71%, while reducing runtime by 91% and reducing memory usage by 61%.

calibration and fitting methods↗

Dynamic Learning of Correlation Potentials for a Time-Dependent Kohn-Sham System

We develop methods to learn the correlation potential for a time-dependent Kohn-Sham (TDKS) system in one spatial dimension. We start from a low-dimensional two-electron system for which we can numerically solve the time-dependent Schr¨odinger equation; this yields electron densities suitable for training models of the correlation potential. We frame the learning problem as one of optimizing a least-squares objective subject to the constraint that the dynamics obey the TDKS equation. Applying adjoints, we develop efficient methods to compute gradients and thereby learn models of the correlation potential. Our results show that it is possible to learn values of the correlation potential such that the resulting electron densities match ground truth densities. We also show how to learn correlation potential functionals with memory, demonstrating one such model that yields reasonable results for trajectories outside the training set.

97 MATHEMATICS AND COMPUTING↗

The Development of a Generalized Riser Flow Regime Map Based Upon Higher Moment and Chaotic Statistics Using Electrical Capacitance Volume Tomography (ECVT)

Dynamic analyses have been applied to the temporal signals from an Electro Capacitance Volume Tomography instrument located near mid-height on the riser of an industrial-scale cold-flow circulating fluidized bed to characterize gas-solids flow behavior in the riser. Twelve capacitance electrodes surround the cylindrical riser over a height of 1.3 m. The instrument used a neural network deconvolution algorithm to determine the spatially resolved solids fraction recorded at 52 Hz. Experiments were carried out over a range of gas and solids flows in the transport regime using a Geldart Group B bed material, high density polyethylene with mean particle size of 880 μm. The radial solids distribution was found to vary from one-time step to the next between profiles typical of laminar and turbulent flow. The duration of time spent in each of these flow profiles depended upon the operating regime – dilute, core-annular, or fast fluidized bed. The chaotic structure of the temporal data was characterized using the three conventional approaches: the first 4 moments from the distribution of signal in time, system memory parameters from the autocorrelation function and the Hurst exponent, and analysis of the correlationentropy and correlation dimension of the attractor. These signal analysis techniques were used to clearly distinguish differences between different transport operating regimes. Specifically, it was experimentally observed that a riser transitions from core annular flow profile to dilute and dense regimes via increasing the frequency of short term transients to either dilute or dense flow profiles, respectively. A regime map was generated based upon these dynamics using solids flux and gas velocity axes. Fast fluidized, core annular, and dilute each exhibited different degree of dynamic characteristics typical of fluid dominated or particle compromising behavior. It should be noted that the magnitude for the different statistics was in the same range regardless of the regime, it was the radial profile for the statistic that changed and subsequently identified that there was a change in the regime. Finally, a reduced regime map was developed consisting of plotting the gas velocity normalized by the upper transport velocity versus the solids flux normalized by the saturation carrying capacity. The use of this reduced plot allowed the data from widely different conditions to be plotted and compared on the same<p>graph. Note that in many instances, some of the statistics identified the operating point as being in one regime while others indicated that it was in another indicating a transition region between dilute or core annular regimes and between the core annular and fast fluidization regimes. This now provides a tool that can be used to optimize process performance, identify changes in operating states, or replicate process dynamics during process scaling or changing operating parameters. </p>

Breault, Ronald↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Hierarchical and Parallelizable Direct Volume Rendering for Irregular and Multiple Grids

A general volume rendering technique is described that efficiently produces images of excellent quality from data defined over irregular grids having a wide variety of formats. Rendering is done in software, eliminating the need for special graphics hardware, as well as any artifacts associated with graphics hardware. Images of volumes with about one million cells can be produced in one to several minutes on a workstation with a 150 MHz processor. A significant advantage of this method for applications such as computational fluid dynamics is that it can process multiple intersecting grids. Such grids present problems for most current volume rendering techniques. Also, the wide range of cell sizes (by a factor of 10,000 or more), which is typical of such applications, does not present difficulties, as it does for many techniques. A spatial hierarchical organization makes it possible to access data from a restricted region efficiently. The tree has greater depth in regions of greater detail, determined by the number of cells in the region. It also makes it possible to render useful 'preview' images very quickly (about one second for one-million-cell grids) by displaying each region associated with a tree node as one cell. Previews show enough detail to navigate effectively in very large data sets. The algorithmic techniques include use of a kappa-d tree, with prefix-order partitioning of triangles, to reduce the number of primitives that must be processed for one rendering, coarse-grain parallelism for a shared-memory MIMD architecture, a new perspective transformation that achieves greater numerical accuracy, and a scanline algorithm with depth sorting and a new clipping technique.

Wilhelms, Jane↗

Stabilizing Non-Abelian Topological Order Against Heralded Noise via Local Lindbladian Dynamics

An important open question for the current generation of highly controllable quantum devices is understanding which phases can be realized as stable steady states under local quantum dynamics. In this work, we show how robust steady-state phases with both Abelian and non-Abelian mixed-state topological order can be stabilized, in two spatial dimensions, against generic “heralded” noise using active dynamics that incorporate measurement and feedback, modeled as a fully local Lindblad master equation. These topologically ordered steady states are two-way connected to pure topologically ordered ground states using local quantum channels, and preserve quantum information for a time that is exponentially large in the system size. Specifically, we present explicit constructions of families of local Lindbladians for both Abelian (ℤ 2 ) and non-Abelian (𝐷 4 ) topological order whose steady states host mixed-state topological order when the noise is below a threshold strength. As the noise strength is increased, these models exhibit first-order transitions to intermediate mixed-state phases where they encode robust classical memories, followed by (first-order) transitions to a trivial steady state at high noise rates. When the noise is imperfectly heralded, steady-state order disappears but our active dynamics significantly enhances the lifetime of the encoded logical information. To carry out the numerical simulations for the non-Abelian 𝐷 4 case, we introduce a generalized stabilizer tableau formalism that permits efficient simulation of the non-Abelian Lindbladian dynamics.

Monte Carlo methods↗

Development of the tangent linear and adjoint models of the global online chemical transport model MPAS-CO 2 v7.3

We describe the development of the tangent linear (TL) and adjoint models of the Model for Prediction Across Scales (MPAS)-CO 2 transport model, which is a global online chemical transport model developed upon the non-hydrostatic Model for Prediction Across Scales – Atmosphere (MPAS-A). The primary goal is to make the model system a valuable research tool for investigating atmospheric carbon transport and inverse modeling. First, we develop the TL code, encompassing all CO 2 transport processes within the MPAS-CO 2 forward model. Then, we construct the adjoint model using a combined strategy involving re-calculation and storage of the essential meteorological variables needed for CO 2 transport. This strategy allows the adjoint model to undertake a long-period integration with moderate memory demands. To ensure accuracy, the TL and adjoint models undergo vigorous verifications through a series of standard tests. The adjoint model, through backward-in-time integration, calculates the sensitivity of atmospheric CO 2 observations to surface CO 2 fluxes and the initial atmospheric CO 2 mixing ratio. To demonstrate the utility of the newly developed adjoint model, we conduct simulations for two types of atmospheric CO 2 observations, namely the tower-based in situ CO 2 mixing ratio and satellite-derived column-averaged CO 2 mixing ratio (X CO 2 ). A comparison between the sensitivity to surface flux calculated by the MPAS-CO 2 adjoint model with its counterpart from CarbonTracker–Lagrange (CT-L) reveals a spatial agreement but notable magnitude differences. These differences, particularly evident for X CO 2 , might be attributed to the two model systems' differences in the simulation configuration, spatial resolution, and treatment of vertical mixing processes. Moreover, this comparison highlights the substantial loss of information in the atmospheric CO 2 observations due to CT-L's spatial domain limitation. Furthermore, the adjoint sensitivity analysis demonstrates that the sensitivities to both surface flux and initial CO 2 conditions spread out throughout the entire Northern Hemisphere within a month. MPAS-CO 2 forward, TL, and adjoint models stand out for their calculation efficiency and variable-resolution capability, making them competitive in computational cost. In conclusion, the successful development of the MPAS-CO 2 TL and adjoint models, and their integration into the MPAS-CO 2 system, establish the possibility of using MPAS's unique features in atmospheric CO 2 transport sensitivity studies and in inverse modeling with advanced methods such as variational data assimilation.

54 ENVIRONMENTAL SCIENCES↗

Single-Chip FPGA Azimuth Pre-Filter for SAR

A field-programmable gate array (FPGA) on a single lightweight, low-power integrated-circuit chip has been developed to implement an azimuth pre-filter (AzPF) for a synthetic-aperture radar (SAR) system. The AzPF is needed to enable more efficient use of data-transmission and data-processing resources: In broad terms, the AzPF reduces the volume of SAR data by effectively reducing the azimuth resolution, without loss of range resolution, during times when end users are willing to accept lower azimuth resolution as the price of rapid access to SAR imagery. The data-reduction factor is selectable at a decimation factor, M, of 2, 4, 8, 16, or 32 so that users can trade resolution against processing and transmission delays. In principle, azimuth filtering could be performed in the frequency domain by use of fast-Fourier-transform processors. However, in the AzPF, azimuth filtering is performed in the time domain by use of finite-impulse-response filters. The reason for choosing the time-domain approach over the frequency-domain approach is that the time-domain approach demands less memory and a lower memory-access rate. The AzPF operates on the raw digitized SAR data. The AzPF includes a digital in-phase/quadrature (I/Q) demodulator. In general, an I/Q demodulator effects a complex down-conversion of its input signal followed by low-pass filtering, which eliminates undesired sidebands. In the AzPF case, the I/Q demodulator takes offset video range echo data to the complex baseband domain, ensuring preservation of signal phase through the azimuth pre-filtering process. In general, in an SAR I/Q demodulator, the intermediate frequency (fI) is chosen to be a quarter of the range-sampling frequency and the pulse-repetition frequency (fPR) is chosen to be a multiple of fI. The AzPF also includes a polyphase spatial-domain pre-filter comprising four weighted integrate-and-dump filters with programmable decimation factors and overlapping phases. To prevent aliasing of signals, the bandwidth of the AzPF is made 80 percent of fPR/M. The choice of four as the number of overlapping phases is justified by prior research in which it was shown that a filter of length 4M can effect an acceptable transfer function. The figure depicts prototype hardware comprising the AzPF and ancillary electronic circuits. The hardware was found to satisfy performance requirements in real-time tests at a sampling rate of 100 MHz.

Gudim, Mimi↗

Current capabilities for simulating the extreme distortion of thin structures subjected to severe impacts

The explicit transient dynamics technology in use today for simulating the impact and subsequent transient dynamic response of a structure has its origins in the 'hydrocodes' dating back to the late 1940's. The growth in capability in explicit transient dynamics technology parallels the growth in speed and size of digital computers. Computer software for simulating the explicit transient dynamic response of a structure is characterized by algorithms that use a large number of small steps. In explicit transient dynamics software there is a significant emphasis on speed and simplicity. The finite element technology used to generate the spatial discretization of a structure is based on a compromise between completeness of the representation for the physical processes modelled and speed in execution. That is, since it is expected in every calculation that the deformation will be finite and the material will be strained beyond the elastic range, the geometry and the associated gradient operators must be reconstructed, as well as complex stress-strain models evaluated at every time step. As a result, finite elements derived for explicit transient dynamics software use the simplest and barest constructions possible for computational efficiency while retaining an essential representation of the physical behavior. The best example of this technology is the four-node bending quadrilateral derived by Belytschko, Lin and Tsay. Today, the speed, memory capacity and availability of computer hardware allows a number of the previously used algorithms to be 'improved.' That is, it is possible with today's computing hardware to modify many of the standard algorithms to improve their representation of the physical process at the expense of added complexity and computational effort. The purpose is to review a number of these algorithms and identify the improvements possible. In many instances, both the older, faster version of the algorithm and the improved and somewhat slower version of the algorithm are found implemented together in software. Specifically, the following seven algorithmic items are examined: the invariant time derivatives of stress used in material models expressed in rate form; incremental objectivity and strain used in the numerical integration of the material models; the use of one-point element integration versus mean quadrature; shell elements used to represent the behavior of thin structural components; beam elements based on stress-resultant plasticity versus cross-section integration; the fidelity of elastic-plastic material models in their representation of ductile metals; and the use of Courant subcycling to reduce computational effort.

Key, Samuel W.↗

Localized strain profile in surface electrode array for programmable composite multiferroic devices

In this work, we investigate localized in-plane strains on the microscale, induced by arrays of biased surface electrodes patterned on piezoelectrics. Particular focus is given to the influence that adjacent electrode pairs have on one another to study the impact of densely packed electrode arrays. We present a series of X-ray microdiffraction studies to reveal the spatially resolved micrometer-scale strain distribution. The strain maps with micrometer-scale resolution highlight how the local strain profile in square regions up to 250 x 250 lm 2 in size is affected by the surface electrodes that are patterned on ferroelectric single-crystal [Pb(Mg 1/3 Nb 2/3 )O 3 ] x -[PbTiO 3 ] 1-x . The experimental measurements and simulation results show the influence of electrode pair distance, positioning of the electrode pair, including the angle of placement, and neighboring electrode pair arrangements on the strength and direction of the regional strain. Our findings are relevant to the development of microarchitected strain-mediated multiferroic devices. The electrode arrays could provide array-addressable localized strain control for applications including straintronic memory, probabilistic computing platforms, microwave devices, and magnetic-activated cell sorting platforms.

42 ENGINEERING↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 km or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed-shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗