Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

A design approach to real-time formatting of high speed multispectral image data

A design approach to formatting multispectral image data in real time at very high data rates is presented for future onboard processing applications. The approach employs a microprocessor-based alternating buffer memory configuration whose formatting function is completely programmable. Data are read from an output buffer in the desired format by applying the proper sequence of addresses to the buffer via a lookup table memory. Sensor data can be processed using this approach at rates limited by the buffer memory access time and the buffer switching process delay time. This design offers flexible high speed data processing and benefits from continuing increases in the performance of digital memories.

Meredith, B. D.↗

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)↗

Intrinsic optical bistability of photon avalanching nanocrystals

Optically bistable materials respond to a single input with two possible optical outputs, contingent on excitation history. Such materials would be ideal for optical switching and memory, but the limited understanding of intrinsic optical bistability (IOB) prevents the development of nanoscale IOB materials suitable for devices. Here we demonstrate IOB in Nd 3+ -doped KPb 2 Cl 5 avalanching nanoparticles, which switch with high contrast between luminescent and non-luminescent states, with hysteresis characteristic of bistability. Here we elucidate a non-thermal mechanism in which IOB originates from suppressed non-radiative relaxation in Nd 3+ ions and from the positive feedback of photon avalanching, resulting in extreme, >200th-order optical nonlinearities. The modulation of laser pulsing tunes the hysteresis widths, and dual-laser excitation enables transistor-like optical switching. This control over nanoscale IOB establishes avalanching nanoparticles for photonic devices in which light is used to manipulate light.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Probing Postmeasurement Entanglement without Postselection

We study the problem of observing quantum collective phenomena emerging from large numbers of measurements. These phenomena are difficult to observe in conventional experiments because, in order to distinguish the effects of measurement from dephasing, it is necessary to postselect on sets of measurement outcomes with Born probabilities that are exponentially small in the number of measurements performed. An unconventional approach, which avoids this exponential “postselection problem”, is to construct cross-correlations between experimental data and the results of simulations on classical computers. However, these cross-correlations generally have no definite relation to physical quantities. We first show how to incorporate classical shadows into this framework, thereby allowing for the construction of quantum information-theoretic cross-correlations. We then identify cross-correlations that both upper and lower bound the measurement-averaged von Neumann entanglement entropy, as well as cross-correlations that lower bound the measurement-averaged purity and entanglement negativity. These bounds show that experiments can be performed to constrain postmeasurement entanglement without the need for postselection. To illustrate our technique, we consider how it could be used to observe the measurement-induced entanglement transition in Haar-random quantum circuits. We use exact numerical calculations as proxies for quantum simulations and, to highlight the fundamental limitations of classical memory, we construct cross-correlations with tensor-network calculations at finite bond dimension. Our results reveal a signature of measurement-induced criticality that can be observed using a quantum simulator in polynomial time and with polynomial classical memory. Published by the American Physical Society 2024

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

Digital processing of satellite imagery application to jungle areas of Peru

The author has identified the following significant results. The use of clustering methods permits the development of relatively fast classification algorithms that could be implemented in an inexpensive computer system with limited amount of memory. Analysis of CCTs using these techniques can provide a great deal of detail permitting the use of the maximum resolution of LANDSAT imagery. Potential cases were detected in which the use of other techniques for classification using a Gaussian approximation for the distribution functions can be used with advantage. For jungle areas, channels 5 and 7 can provide enough information to delineate drainage patterns, swamp and wet areas, and make a reasonable broad classification of forest types.

Pomalaza, J. C.↗

The Lick Observatory charge-coupled device /CCD/ and controller

A description is given of a flexible microprocessor-based controller for charge-coupled device two-dimensional detectors. It is noted that the controller can operate under manual control or as a slave to a remote computer linked by coaxial cable. The system is discussed, together with data taken at the telescope and in the laboratory. The flexibility of the controller derives from its modular form and from the use of a programmable 8-bit microprocessor to control and sequence the electronic logic. The electronic circuits for such functions as signal processing, clock sequencing, voltage level adjustment, and temperature control are on small individual plug-in cards, making future improvements and changes simple. The software of the microprocessor is stored in erasable, programmable, read-only memory. Among the limitations of the controller is a scan speed of roughly 35 microsec per pixel.

Robinson, L. B.↗

KAM (Knowledge Acquisition Module): A tool to simplify the knowledge acquisition process

Analysts, knowledge engineers and information specialists are faced with increasing volumes of time-sensitive data in text form, either as free text or highly structured text records. Rapid access to the relevant data in these sources is essential. However, due to the volume and organization of the contents, and limitations of human memory and association, frequently: (1) important information is not located in time; (2) reams of irrelevant data are searched; and (3) interesting or critical associations are missed due to physical or temporal gaps involved in working with large files. The Knowledge Acquisition Module (KAM) is a microcomputer-based expert system designed to assist knowledge engineers, analysts, and other specialists in extracting useful knowledge from large volumes of digitized text and text-based files. KAM formulates non-explicit, ambiguous, or vague relations, rules, and facts into a manageable and consistent formal code. A library of system rules or heuristics is maintained to control the extraction of rules, relations, assertions, and other patterns from the text. These heuristics can be added, deleted or customized by the user. The user can further control the extraction process with optional topic specifications. This allows the user to cluster extracts based on specific topics. Because KAM formalizes diverse knowledge, it can be used by a variety of expert systems and automated reasoning applications. KAM can also perform important roles in computer-assisted training and skill development. Current research efforts include the applicability of neural networks to aid in the extraction process and the conversion of these extracts into standard formats.

Gettig, Gary A.↗

Proposed data compression schemes for the Galileo S-band contingency mission

The Galileo spacecraft is currently on its way to Jupiter and its moons. In April 1991, the high gain antenna (HGA) failed to deploy as commanded. In case the current efforts to deploy the HGA fails, communications during the Jupiter encounters will be through one of two low gain antenna (LGA) on an S-band (2.3 GHz) carrier. A lot of effort has been and will be conducted to attempt to open the HGA. Also various options for improving Galileo's telemetry downlink performance are being evaluated in the event that the HGA will not open at Jupiter arrival. Among all viable options the most promising and powerful one is to perform image and non-image data compression in software onboard the spacecraft. This involves in-flight re-programming of the existing flight software of Galileo's Command and Data Subsystem processors and Attitude and Articulation Control System (AACS) processor, which have very limited computational and memory resources. In this article we describe the proposed data compression algorithms and give their respective compression performance. The planned image compression algorithm is a 4 x 4 or an 8 x 8 multiplication-free integer cosine transform (ICT) scheme, which can be viewed as an integer approximation of the popular discrete cosine transform (DCT) scheme. The implementation complexity of the ICT schemes is much lower than the DCT-based schemes, yet the performances of the two algorithms are indistinguishable. The proposed non-image compression algorith is a Lempel-Ziv-Welch (LZW) variant, which is a lossless universal compression algorithm based on a dynamic dictionary lookup table. We developed a simple and efficient hashing function to perform the string search.

Cheung, Kar-Ming↗

Improved reduced-resolution satellite imagery

The resolution of satellite imagery is often traded-off to satisfy transmission time and bandwidth, memory, and display limitations. Although there are many ways to achieve the same reduction in resolution, algorithms vary in their ability to preserve the visual quality of the original imagery. These issues are investigated in the context of the Landsat browse system, which permits the user to preview a reduced resolution version of a Landsat image. Wavelets-based techniques for resolution reduction are proposed as alternatives to subsampling used in the current system. Experts judged imagery generated by the wavelets-based methods visually superior, confirming initial quantitative results. In particular, compared to subsampling, the wavelets-based techniques were much less likely to obscure roads, transmission lines, and other linear features present in the original image, introduce artifacts and noise, and otherwise reduce the usefulness of the image. The wavelets-based techniques afford multiple levels of resolution reduction and computational speed. This study is applicable to a wide range of reduced resolution applications in satellite imaging systems, including low resolution display, spaceborne browse, emergency image transmission, and real-time video downlinking.

Ellison, James↗

Comparison of the NASCAP/GEO, POLAR, SEE Charging Handbook, and NASCAP-2K.1 Spacecraft Charging Codes

The NASA Charging Analyzer Program (NASCAP) spacecraft charging software developed by Maxwell Technologies has been widely used for the past fifteen to twenty years in satellite design and investigation of spacecraft charging related anomalies. Individual versions of the NASCAP software are available for use in low inclination, low Earth orbit environments (NASCAP[LEO) and geostationary orbit environments (NASCAP/GEO). In addition, the Potentials of Large objects in the Auroral Region (POLAR) code is available for use in LEO polar orbit environments. NASCAP/GEO and POLAR were both written in the 1980's using algorithms appropriate for the computers of the time. They solve the Poisson-Vlasov system for currents and densities assuming limited speed and memory of computer systems standard for the day. In addition, use of the charging models requires individual input files that are not readily transported into the various codes to facilitate comparison of results by the user.

Neergaard, Linda E.↗

Evaluation Metrics for the Paragon XP/S-15

On February 17th 1993, the Numerical Aerodynamic Simulation (NAS) facility located at the NASA Ames Research Center installed a 224 node Intel Paragon XP/S-15 system. After its installation, the Paragon was found to be in a very immature state and was unable to support a NAS users' workload, composed of a wide range of development and production activities. As a first step towards addressing this problem, we implemented a set of metrics to objectively monitor the system as operating system and hardware upgrades were installed. The metrics were designed to measure four aspects of the system that we consider essential to support our workload: availability, utilization, functionality, and performance. This report presents the metrics collected from February 1993 to August 1993. Since its installation, the Paragon availability has improved from a low of 15% uptime to a high of 80%, while its utilization has remained low. Functionality and performance have improved from merely running one of the NAS Parallel Benchmarks to running all of them faster (between 1 and 2 times) than on the iPSC/860. In spite of the progress accomplished, fundamental limitations of the Paragon operating system are restricting the Paragon from supporting the NAS workload. The maximum operating system message passing (NORMA IPC) bandwidth was measured at 11 Mbytes/s, well below the peak hardware bandwidth (175 Mbytes/s), limiting overall virtual memory and Unix services (i.e. Disk and HiPPI I/O) performance. The high NX application message passing latency (184 microns), three times than on the iPSC/860, was found to significantly degrade performance of applications relying on small message sizes. The amount of memory available for an application was found to be approximately 10 Mbytes per node, indicating that the OS is taking more space than anticipated (6 Mbytes per node).

Traversat, Bernard↗

Parallel Rendering of Large Time-Varying Volume Data

Interactive visualization of large time-varying 3D volume datasets has been and still is a great challenge to the modem computational world. It stretches the limits of the memory capacity, the disk space, the network bandwidth and the CPU speed of a conventional computer. In this SURF project, we propose to develop a parallel volume rendering program on SGI's Prism, a cluster computer equipped with state-of-the-art graphic hardware. The proposed program combines both parallel computing and hardware rendering in order to achieve an interactive rendering rate. We use 3D texture mapping and a hardware shader to implement 3D volume rendering on each workstation. We use SGI's VisServer to enable remote rendering using Prism's graphic hardware. And last, we will integrate this new program with ParVox, a parallel distributed visualization system developed at JPL. At the end of the project, we Will demonstrate remote interactive visualization using this new hardware volume renderer on JPL's Prism System using a time-varying dataset from selected JPL applications.

Garbutt, Alexander E.↗

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering↗

Numerical Prediction Methods (Reynolds-Averaged Navier-Stokes Simulations of Transonic Separated Flows)

During the past five years, numerous pioneering archival publications have appeared that have presented computer solutions of the mass-weighted, time-averaged Navier-Stokes equations for transonic problems pertinent to the aircraft industry. These solutions have been pathfinders of developments that could evolve into a major new technological capability, namely the computational Navier-Stokes technology, for the aircraft industry. So far these simulations have demonstrated that computational techniques, and computer capabilities have advanced to the point where it is possible to solve forms of the Navier-Stokes equations for transonic research problems. At present there are two major shortcomings of the technology: limited computer speed and memory, and difficulties in turbulence modelling and in computation of complex three-dimensional geometries. These limitations and difficulties are the pacing items of the continuing developments, although the one item that will most likely turn out to be the most crucial to the progress of this technology is turbulence modelling. The objective of this presentation is to discuss the state of the art of this technology and suggest possible future areas of research. We now discuss some of the flow conditions for which the Navier-Stokes equations appear to be required. On an airfoil there are four different types of interaction of a shock wave with a boundary layer: (1) shock-boundary-layer interaction with no separation, (2) shock-induced turbulent separation with immediate reattachment (we refer to this as a shock-induced separation bubble), (3) shock-induced turbulent separation without reattachment, and (4) shock-induced separation bubble with trailing edge separation.

Mehta, Unmeel↗

Electrically Variable Resistive Memory Devices

Nonvolatile electronic memory devices that store data in the form of electrical- resistance values, and memory circuits based on such devices, have been invented. These devices and circuits exploit an electrically-variable-resistance phenomenon that occurs in thin films of certain oxides that exhibit the colossal magnetoresistive (CMR) effect. It is worth emphasizing that, as stated in the immediately preceding article, these devices function at room temperature and do not depend on externally applied magnetic fields. A device of this type is basically a thin film resistor: it consists of a thin film of a CMR material located between, and in contact with, two electrical conductors. The application of a short-duration, low-voltage current pulse via the terminals changes the electrical resistance of the film. The amount of the change in resistance depends on the size of the pulse. The direction of change (increase or decrease of resistance) depends on the polarity of the pulse. Hence, a datum can be written (or a prior datum overwritten) in the memory device by applying a pulse of size and polarity tailored to set the resistance at a value that represents a specific numerical value. To read the datum, one applies a smaller pulse - one that is large enough to enable accurate measurement of resistance, but small enough so as not to change the resistance. In writing, the resistance can be set to any value within the dynamic range of the CMR film. Typically, the value would be one of several discrete resistance values that represent logic levels or digits. Because the number of levels can exceed 2, a memory device of this type is not limited to binary data. Like other memory devices, devices of this type can be incorporated into a memory integrated circuit by laying them out on a substrate in rows and columns, along with row and column conductors for electrically addressing them individually or collectively.

Liu, Shangqing↗