Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “contiguous memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Threaded Multi-Core GEMM with MoA and Cache-Blocking: Preprint

A threaded multi-core implementation of the high performance dense linear algebra matrix-matrix multiply GEMM kernel is described. This kernel is widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses by employing outer-product forms. Our performance studies demonstrate that the MoA implementation of double precision DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on the Intel Xeon Skylake processor over the vendor supplied Intel MKL basic linear algebra libraries. Results are presented for the NREL Eagle supercomputer. The multi-core DGEMM achieves over 100 GigaFlops/sec with eight openMP threads.

cache-blocking↗

Improving the Performance of DGEMM with MoA and Cache-Blocking: Preprint

The goal of this paper is to demonstrate performance enhancements of the high performance dense linear algebra matrix-matrix multiply DGEMM kernel, widely implemented by vendors in the basic linear algebra subroutine BLAS library. The mathematics of arrays (MoA) paradigm due to Mullin (1988) results in contiguous memory accesses in combination with Church-Rosser complete language constructs optimized for target processor architectures [3]. Our performance studies demonstrate that the MoA implementation of DGEMM combined with optimal cache-blocking strategies results in at least a 25% performance gain on both Intel Xeon Skylake and IBM Power-9 processors over the vendor supplied Intel MKL and IBM ESSL basic linear algebra libraries. Results are presented for the NREL Eagle and ORNL Summit supercomputers.

cache-blocking↗

MATAR: A performance portability and productivity implementation of data-oriented design with Kokkos

There is a need for simple, fast, and memory-efficient multidimensional data structures for dense and sparse storage that arise with numerical methods and in software applications. The data structures must perform equally well across multiple computer architectures, including CPUs and GPUs. For this purpose, we developed MATAR, a C++ software library that allows for simple creation and use of intricate data structures that is also portable across disparate architectures using Kokkos. Here, the performance aspect is achieved by forcing contiguous memory layout (or as close to contiguous as possible) for multidimensional and multi-size dense or sparse MATrix and ARray (hence, MATAR) types. Our results show that MATAR has the capability to improve memory utilization, performance, and programmer productivity in scientific computing. This is achieved by fitting more work into the available memory, minimizing memory loads required, and by loading memory in the most efficient order. This document describes the purpose of the work, the implementation of each of the data types, and the resulting performance both in some simple baseline test cases and in an application code.

97 MATHEMATICS AND COMPUTING↗

Out-of-Core Streamline Visualization on Large Unstructured Meshes

It's advantageous for computational scientists to have the capability to perform interactive visualization on their desktop workstations. For data on large unstructured meshes, this capability is not generally available. In particular, particle tracing on unstructured grids can result in a high percentage of non-contiguous memory accesses and therefore may perform very poorly with virtual memory paging schemes. The alternative of visualizing a lower resolution of the data degrades the original high-resolution calculations. This paper presents an out-of-core approach for interactive streamline construction on large unstructured tetrahedral meshes containing millions of elements. The out-of-core algorithm uses an octree to partition and restructure the raw data into subsets stored into disk files for fast data retrieval. A memory management policy tailored to the streamline calculations is used such that during the streamline construction only a very small amount of data are brought into the main memory on demand. By carefully scheduling computation and data fetching, the overhead of reading data from the disk is significantly reduced and good memory performance results. This out-of-core algorithm makes possible interactive streamline visualization of large unstructured-grid data sets on a single mid-range workstation with relatively low main-memory capacity: 5-20 megabytes. Our test results also show that this approach is much more efficient than relying on virtual memory and operating system's paging algorithms.

Ueng, Shyh-Kuang↗

Microwave Imaging on Metal Objects

This final report for the project discusses the attempts to model, using different methods, microwave image reconstruction. Maximum Entropy Method was not successful. Attempts to use Singular Value Decomposition (SVD) got some good results after initial failure. SVD is based upon a theory of linear algebra, to the effect that any M X N Matrix A whose number of rows M is greater than or equal to its number of columns, N can be written as the product of an M X N column-orthogonal matrix U, an N X N diagonal Matrix, W, with m positive or zero elements (the singular values) and the transposition of an N X N orthogonal matrix V. In microwave imaging, the scattered fields can be expressed by the induced current distribution. The SVD method required more contiguous computer memory than was available. Work was also done on the Conjugate Gradient Method (CGM), which didn't work well when tried earlier. It was found that separation of the imaginary part and the real part during calculation may work. This work was considered incomplete as of the end of the grant period.

Tolliver, C. L.↗

Atmospheric Structure Prediction for Infrasound Propagation Modeling Using Deep Learning

Abstract Infrasound is generated by a variety of natural and anthropogenic sources. Infrasonic waves travel through the dynamic atmosphere, which can change on the order of minutes to hours. Infrasound propagation largely depends on the wind and temperature structure of the atmosphere. Numerical weather prediction models are available to provide atmospheric specifications, but uncertainties in these models exist and they are computationally expensive to run. Machine learning has proven useful in predicting tropospheric weather using Long Short‐Term Memory (LSTM) networks. An LSTM network is utilized to make atmospheric specification predictions up to ∼30 km for three different training and testing scenarios: (a) the model is trained and tested using only radiosonde data from the Albuquerque, NM, USA station, (b) the model is trained on radiosonde stations across the contiguous US, excluding the Albuquerque, NM, USA station, which was reserved for testing, and (c) the model is trained and tested on radiosonde stations across the contiguous US. Long Short‐Term Memory predictions are compared to a state‐of‐the‐art reanalysis model and show cases where the LSTM outperforms, performs equally as well, or underperforms in comparison to the state‐of‐the‐art. Regional and temporal trends in model performance across the US are also discussed. Results suggest that the LSTM model is a viable tool for predicting atmospheric specifications for infrasound propagation modeling.

54 ENVIRONMENTAL SCIENCES↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Techniques for storing data to enhance recovery and detection of data corruption errors

Often there are errors when reading data from computer memory. To detect and correct these errors, there are multiple types of error correction codes. Disclosed is an error correction architecture that creates a codeword having a data portion and an error correction code portion. Swizzling rearranges the order of bits and distributes the bits among different codewords. Because the data is redistributed, a potential memory error of up to N contiguous bits, where N for example equals 2 times the number of codewords swizzled together, only affects up to, at most, two bits per swizzled codeword. This keeps the error within the error detecting capabilities of the error correction architecture. Furthermore, this can allow improved error correction and detection without requiring a change to error correcting code generators and checkers.

Mills, Peter↗

Performance analysis of replication ALOHA for fading mobile communications channels

This paper describes an ALOHA random access protocol for fading communications channels. A two-state Markov model is used for the channel error process to account for the channel fading memory. The ALOHA protocol is modified to send multiple contiguous copies of a message at each transmission attempt. Both pure and slotted ALOHA channels are considered. The analysis is applicable to fading environments where the channel memory is short compared to the propagation delay. It is shown that smaller delay may be achieved using replications and, in noisy conditions, can also improve throughput.

Yan, Tsun-Yee↗

Increasing phosphorus loss despite widespread concentration decline in US rivers

The loss of phosphorous (P) from the land to aquatic systems has polluted waters and threatened food production worldwide. Systematic trend analysis of P, a nonrenewable resource, has been challenging, primarily due to sparse and inconsistent historical data. Here, we leveraged intensive hydrometeorological data and the recent renaissance of deep learning approaches to fill data gaps and reconstruct temporal trends. We trained a multitask long short-term memory model for total P (TP) using data from 430 rivers across the contiguous United States (CONUS). Trend analysis of reconstructed daily records (1980–2019) shows widespread decline in concentrations, with declining, increasing, and insignificantly changing trends in 60%, 28%, and 12% of the rivers, respectively. Concentrations in urban rivers have declined the most despite rising urban population in the past decades; concentrations in agricultural rivers however have mostly increased, suggesting not-as-effective controls of nonpoint sources in agriculture lands compared to point sources in cities. TP loss, calculated as fluxes by multiplying concentration and discharge, however exhibited an overall increasing rate of 6.5% per decade at the CONUS scale over the past 40 y, largely due to increasing river discharge. Results highlight the challenge of reducing TP loss that is complicated by changing river discharge in a warming climate.

Science & Technology - Other Topics↗

Temperature outweighs light and flow as the predominant driver of dissolved oxygen in US rivers

The concentration of dissolved oxygen (DO), an important measure of water quality and river metabolism, varies tremendously in time and space. Riverine DO is commonly perceived as regulated by interacting and competing drivers (light, temperature and flow) that define rivers’ climate. Its continental-scale drivers, however, have remained elusive, partly due to the scarcity and spatio-temporal inconsistency of water quality data. Here we show, via a deep learning model (long short-term memory) trained using data from 580 rivers, that temperature predominantly drives daily DO dynamics in the contiguous United States. Light comes a close second, whereas flow imparts minimal influence. This work showcases the promise of using deep learning models for data filling that enables large-scale systematic analysis of patterns and drivers. Results show fairly accurate prediction of DO by temperature alone, and declining DO in warming rivers, which has important implications for water security and ecosystem health in the future climate.

54 ENVIRONMENTAL SCIENCES↗

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

St. Joseph Peninsula Disasters: Using NASA Earth Observations to Investigate Land Cover, Shoreline Change, and Sediment Transport in St. Joseph Peninsula after Hurricane Michael

T.H. Stone Memorial St. Joseph Peninsula State Park experienced significant damages from Hurricane Michael in 2018, the first Category 5 hurricane to hit the contiguous United States since 1992. These damages included a 300-meter-wide and 10-meter-deep breach in the peninsula, habitat disruption, and a forced closure of over half of the total park area. These damages, coupled with restricted visitor access, resulted in a significant loss of revenue for the park. NASA DEVELOP partnered with the Florida Department of Environmental Protection (DEP) to determine the overall impact of Hurricane Michael on land cover and shoreline change by using NASA Earth observations including Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI), Aqua Moderate Resolution Imaging Spectroradiometer (MODIS), and the European Space Agency’s Sentinel-2 Multispectral Instrument (MSI) to analyze sediment transport and climatology to further understand the lasting impacts of hurricanes on the ecosystems of the park. The DEVELOP team’s analyses showed that chlorophyll-a concentrations, sea surface temperature, and precipitation are increasing over time. The sediment transport analysis showed dynamic movement across the peninsula, with the greatest erosion occurring within the bay and along the length of the peninsula. These results are supported by evidence of declining seagrass abundances and seasonal turbidity patterns within those areas. Providing these analyses for the partner allows for a greater understanding of how best to proceed with restoration efforts, which may include rebuilding camping services, expanding fishing recreation, and conserving habitats for endangered species.

Erica Kriner↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

The SGI/Cray T3E: Experiences and Insights

The NASA Goddard Space Flight Center is home to the fifth most powerful supercomputer in the world, a 1024 processor SGI/Cray T3E-600. The original 512 processor system was placed at Goddard in March, 1997 as part of a cooperative agreement between the High Performance Computing and Communications Program's Earth and Space Sciences Project (ESS) and SGI/Cray Research. The goal of this system is to facilitate achievement of the Project milestones of 10, 50 and 100 GFLOPS sustained performance on selected Earth and space science application codes. The additional 512 processors were purchased in March, 1998 by the NASA Earth Science Enterprise for the NASA Seasonal to Interannual Prediction Project (NSIPP). These two "halves" still operate as a single system, and must satisfy the unique requirements of both aforementioned groups, as well as guest researchers from the Earth, space, microgravity, manned space flight and aeronautics communities. Few large scalable parallel systems are configured for capability computing, so models are hard to find. This unique environment has created a challenging system administration task, and has yielded some insights into the supercomputing needs of the various NASA Enterprises, as well as insights into the strengths and weaknesses of the T3E architecture and software. The T3E is a distributed memory system in which the processing elements (PE's) are connected by a low latency, high bandwidth bidirectional 3-D torus. Due to the focus on high speed communication between PE's, the T3E requires PE's to be allocated contiguously per job. Further, jobs will only execute on the user specified number of PE's and PE timesharing is possible but impractical. With a highly varied job mix in both size and runtime of jobs, the resulting scenario is PE fragmentation and an inability to achieve near 100% utilization. SGI/Cray has provided several scheduling and configuration tools to minimize the impact of fragmentation. These tools include PScheD (the political scheduler), GRM (the global resource manager) and NQE (the Network Queuing Environment). Features and impact of these tools will be discussed, as will resulting performance and utilization data. As a distributed memory system, the T3E is designed to be programmed through explicit message passing. Consequently, certain assumptions related to code design are made by the operating system (UNICOS/mk) and its scheduling tools. With the exception of HPF, which does run on the T3E, however poorly, alternative programming styles have the potential to impact the T3E in unexpected and undesirable ways. Several examples will be presented (preceeded with the disclaimer, "Don't try this at home! Violators will be prosecuted!")

Bernard, Lisa Hamet↗

Performance analysis of the ALOHA protocol with replication in a fading channel for the Mobile Satellite Experiment

The analysis of the ALOHA random access protocol for communications channels with fading is presented. The protocol is modified to send multiple contiguous copies of a message at each transmission attempt. Both pure and slotted ALOHA channels are considered. A general two state model is used for the channel error process to account for the channel fading memory. It is shown that greater throughput and smaller delay may be achieved using repetitions. The model is applied to the analysis of the delay-throughput performance in a fading mobile communications environment. Numerical results are given for NASA's Mobile Satellite Experiment.

Clare, L. P.↗

Computing rank‐revealing factorizations of matrices stored out‐of‐core

This paper describes efficient algorithms for computing rank-revealing factorizations of matrices that are too large to fit in main memory (RAM), and must instead be stored on slow external memory devices such as disks (out-of-core or out-of-memory). Traditional algorithms for computing rank-revealing factorizations (such as the column pivoted QR factorization and the singular value decomposition) are very communication intensive as they require many vector-vector and matrix-vector operations, which become prohibitively expensive when data is not in RAM. Randomization allows to reformulate new methods so that large contiguous blocks of the matrix are processed in bulk. The paper describes two distinct methods. The first is a blocked version of column pivoted Householder QR, organized as a “left-looking” method to minimize the number of the expensive write operations. The second method results employs a UTV factorization. It is organized as an algorithm-by-blocks to overlap computations and I/O operations. As it incorporates power iterations, it is much better at revealing the numerical rank. Numerical experiments on several computers demonstrate that the new algorithms are almost as fast when processing data stored on slow memory devices as traditional algorithms are for data stored in RAM.

97 MATHEMATICS AND COMPUTING↗

Improving the prediction of daily reservoir releases over the CONUS using conditioned LSTM

Reservoirs play a vital role in regulating streamflow timing and variability for hydroelectricity, flood control, water supply, irrigation, and recreation. Despite their importance, many reservoirs lack comprehensive operational guidelines, making their management complex due to conflicting operational objectives. Hence traditional policy-based reservoir models often fail to capture real-world conditions accurately and they depend on perfect streamflow predictions, which are not always available. In contrast, data-driven models like Long Short-Term Memory (LSTM) networks offer a robust alternative. This study introduces an approach that integrates reservoir characteristics—such as main use, climate, and maximum capacity—into the LSTM model to enhance reservoir release predictions. Using data from nearly 200 reservoirs in the contiguous United States (CONUS), our conditioned LSTM model (LSTM_cond) was compared with both the vanila LSTM and a traditional policy-based approach. Furthermore, our results show that while both LSTM_cond and LSTM perfoms better than the policy-based approach, LSTM_cond consistently outperforms LSTM for hydroelectric, water supply, irrigation, and recreation reservoirs. The KGE median values for LSTM_cond for out-sample reservoirs are 0.764, 0.565, 0.821, and 0.779, respectively, for the aforementioned reservoir types, which are consistently higher that the corresponding KGE values of 0.737, 0.413, 0.775, and 0.713 of LSTM, demonstrating its advantages in improving generalizability.

CONUS↗