Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “masking algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Coronagraph Design Optimization for Segmented Aperture Telescopes

The goal of directly imaging Earth-like planets in the habitable zone of other stars has motivated the design of coronagraphs for use with large segmented aperture space telescopes. In order to achieve an optimal trade-o between planet light throughput and di racted starlight suppression, we consider coronagraphs comprised of a stage of phase control implemented with deformable mirrors (or other optical elements), pupil plane apodization masks (gray scale or complex valued), and focal plane masks (either amplitude only or complex-valued, including phase only such as the vector vortex coronagraph). The optimization of these optical elements, with the goal of achieving 10 or more orders of magnitude in the suppression of on-axis (starlight) di racted light, represents a challenging non-convex optimization problem with a nonlinear dependence on control degrees of freedom. We develop a new algorithmic approach to the design optimization problem, which we call the "Auxiliary Field Optimization" (AFO) algorithm. The central idea of the algorithm is to embed the original optimization problem, for either phase or amplitude (apodization) in various planes of the coronagraph, into a problem containing additional degrees of freedom, speci cally ctitious "auxiliary" electric elds which serve as targets to inform the variation of our phase or amplitude parameters leading to good feasible designs. We present the algorithm, discuss details of its numerical implementation, and prove convergence to local minima of the objective function (here taken to be the intensity of the on-axis source in a "dark hole" region in the science focal plane). Finally, we present results showing application of the algorithm to both unobscured o -axis and obscured on-axis segmented telescope aperture designs. The application of the AFO algorithm to the coronagraph design problem has produced solutions which are capable of directly imaging planets in the habitable zone, provided end-to-end telescope system stability requirements can be met. Ongoing work includes advances of the AFO algorithm reported here to design in additional robustness to a resolved star, and other phase or amplitude aberrations to be encountered in a real segmented aperture space telescope.

Redding, Dave↗

Enhancing molecular design efficiency: Uniting language models and generative networks with genetic algorithms

This study examines the effectiveness of generative models in drug discovery, material science, and polymer science, aiming to overcome constraints associated with traditional inverse design methods relying on heuristic rules. Generative models generate synthetic data resembling real data, enabling deep learning model training without extensive labeled datasets. They prove valuable in creating virtual libraries of molecules for material science and facilitating drug discovery by generating molecules with specific properties. While generative adversarial networks (GANs) are explored for these purposes, mode collapse restricts their efficacy, limiting novel structure variability. To address this, we introduce a masked language model (LM) inspired by natural language processing. Although LMs alone can have inherent limitations, we propose a hybrid architecture combining LMs and GANs to efficiently generate new molecules, demonstrating superior performance over standalone masked LMs, particularly for smaller population sizes. This hybrid LM-GAN architecture enhances efficiency in optimizing properties and generating novel samples.

97 MATHEMATICS AND COMPUTING↗

Expectation-propagation for weak radionuclide identification at radiation portal monitors

We propose a sparsity-promoting Bayesian algorithm capable of identifying radionuclide signatures from weak sources in the presence of a high radiation background. The proposed method is relevant to radiation identification for security applications. In such scenarios, the background typically consists of terrestrial, cosmic, and cosmogenic radiation that may cause false positive responses. We evaluate the new Bayesian approach using gamma-ray data and are able to identify weapons-grade plutonium, masked by naturally-occurring radioactive material (NORM), in a measurement time of a few seconds. We demonstrate this identification capability using organic scintillators (stilbene crystals and EJ-309 liquid scintillators), which do not provide direct, high-resolution, source spectroscopic information. Compared to the EJ-309 detector, the stilbene-based detector exhibits a lower identification error, on average, owing to its better energy resolution. Organic scintillators are used within radiation portal monitors to detect gamma rays emitted from conveyances crossing ports of entry. The described method is therefore applicable to radiation portal monitors deployed in the field and could improve their threat discrimination capability by minimizing “nuisance” alarms produced either by NORM-bearing materials found in shipped cargoes, such as ceramics and fertilizers, or radionuclides in recently treated nuclear medicine patients.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

NASA Tech Briefs, April 2006

The topics covered include: 1) Replaceable Sensor System for Bioreactor Monitoring; 2) Unitary Shaft-Angle and Shaft-Speed Sensor Assemblies; 3) Arrays of Nano Tunnel Junctions as Infrared Image Sensors; 4) Catalytic-Metal/PdO(sub x)/SiC Schottky-Diode Gas Sensors; 5) Compact, Precise Inertial Rotation Sensors for Spacecraft; 6) Universal Controller for Spacecraft Mechanisms; 7) The Flostation - an Immersive Cyberspace System; 8) Algorithm for Aligning an Array of Receiving Radio Antennas; 9) Single-Chip T/R Module for 1.2 GHz; 10) Quantum Entanglement Molecular Absorption Spectrum Simulator; 11) FuzzObserver; 12) Internet Distribution of Spacecraft Telemetry Data; 13) Semi-Automated Identification of Rocks in Images; 14) Pattern-Recognition Algorithm for Locking Laser Frequency; 15) Designing Cure Cycles for Matrix/Fiber Composite Parts; 16) Controlling Herds of Cooperative Robots; 17) Modification of a Limbed Robot to Favor Climbing; 18) Vacuum-Assisted, Constant-Force Exercise Device; 19) Production of Tuber-Inducing Factor; 20) Quantum-Dot Laser for Wavelengths of 1.8 to 2.3 micron; 21) Tunable Filter Made From Three Coupled WGM Resonators; and 22) Dynamic Pupil Masking for Phasing Telescope Mirror Segments.

Source record↗

Ice surface temperature retrieval from AVHRR, ATSR, and passive microwave satellite data: Algorithm development and application

During the first half of our second project year we have accomplished the following: (1) acquired a new AVHRR data set for the Beaufort Sea area spanning an entire year; (2) acquired additional ATSR data for the Arctic and Antarctic now totaling over seven months; (3) refined our AVHRR Arctic and Antarctic ice surface temperature (IST) retrieval algorithm, including work specific to Greenland; (4) developed ATSR retrieval algorithms for the Arctic and Antarctic, including work specific to Greenland; (5) investigated the effects of clouds and the atmosphere on passive microwave 'surface' temperature retrieval algorithms; (6) generated surface temperatures for the Beaufort Sea data set, both from AVHRR and SSM/I; and (7) continued work on compositing GAC data for coverage of the entire Arctic and Antarctic. During the second half of the year we will continue along these same lines, and will undertake a detailed validation study of the AVHRR and ATSR retrievals using LEADEX and the Beaufort Sea year-long data. Cloud masking methods used for the AVHRR will be modified for use with the ATSR. Methods of blending in situ and satellite-derived surface temperature data sets will be investigated.

Key, Jeff↗

Approach trajectory planning system for maximum concealment

A computer-simulation study was undertaken to investigate a maximum concealment guidance technique (pop-up maneuver), which military aircraft may use to capture a glide path from masked, low-altitude flight typical of terrain following/terrain avoidance flight enroute. The guidance system applied to this problem is the Fuel Conservative Guidance System. Previous studies using this system have concentrated on the saving of fuel in basically conventional land and ship-based operations. Because this system is based on energy-management concepts, it also has direct application to the pop-up approach which exploits aircraft performance. Although the algorithm was initially designed to reduce fuel consumption, the commanded deceleration is at its upper limit during the pop-up and, therefore, is a good approximation of a minimum-time solution. Using the model of a powered-lift aircraft, the results of the study demonstrated that guidance commands generated by the system are well within the capability of an automatic flight-control system. Results for several initial approach conditions are presented.

Warner, David N., Jr.↗

Derivation of Shortwave Radiometric Adjustments for SNPP and NOAA-20 VIIRS for the NASA MODIS-VIIRS Continuity Cloud Products

Climate studies, including trend detection and other time series analyses, necessarily require stable, well-characterized and long-term data records. For satellite-based geophysical retrieval datasets, such data records often involve merging the observational records of multiple similar, though not identical, instruments. The National Aeronautics and Space Administration (NASA)cloud mask(CLDMSK) and cloud-top and optical properties(CLDPROP) products are designed to bridge the observational records of the Moderate-resolution Imaging Spectroradiometer (MODIS) onboard NASA’s Aqua satellite and the Visible Infrared Imaging Radiometer Suite (VIIRS) onboard the joint NASA/National Oceanic and Atmospheric Administration(NOAA)Suomi National Polar-orbiting Partnership (SNPP)satellite and NOAA’s new generation of operational polar-orbiting weather platforms(NOAA-20+).Early implementations of the CLDPROP algorithms on Aqua MODIS and SNPP VIIRS suffered from large intersensor biases in cloud optical properties that were traced back to relative radiometric inconsistency in analogous shortwave channels on both imagers, with VIIRS generally observing brighter top-of-atmosphere spectral reflectance than MODIS (e.g., up to 5%brighter in the 0.67μm channel). Radiometric adjustment factors for the SNPP and NOAA-20 VIIRS shortwave channels used in the cloud optical property retrievals are derived from an extensive analysis of the overlapping observational records with Aqua MODIS, specifically for homogenous maritime liquid water cloud scenes for which the viewing/solar geometry of MODIS and VIIRS match. Application of these adjustment factors to the VIIRS L1B prior to ingestion into the CLDMSK and CLDPROP algorithms yields improved intersensor agreement, particularly for cloud optical properties

relative radiometry↗

Broadband Performance of TPF's High-contrast Imaging Testbed: Modeling and Simulations

The broadband performance of the high-contrast imaging testbed (HCIT) at JPL is investigated through optical modeling and simulations. The analytical tool is an optical simulation algorithm developed by combining the HCIT's optical model with a speckle-nulling algorithm that operates directly on coronagraphic images, an algorithm identical to the one currently being used on the HCIT to actively suppress scattered light via a deformable mirror. It is capable of performing full three-dimensional end-to-end near-field diffraction analysis on the HCIT's optical system. By conducting speckle-nulling optimization, we clarify the HCIT's capability and limitations in terms of its broadband contrast performance under various realistic conditions. Considered cases include non-ideal occulting masks, such as a mask with optical density and wavelength dependent parasitic phase-delay errors (i.e., a not band-limited occulting mask) and the one with an optical-density profile corresponding to a measured, non-standard profile, as well as the independently measured phase errors of all optics. Most of the information gathered on the HCIT's optical components through measurement and characterization over the last several years at JPL has been used in this analysis to make the predictions as accurate as possible. The best contrast values predicted so far by our simulations obtainable on the HCIT illuminated with a broadband light having a bandwidth of 80nm and centered at 800nm wavelength are Cm=1.1x10-8 (mean) and C4=4.9x10-8 (at 4(lamda)/D), respectively. In this paper we report our preliminary findings about the broadband light performance of the HCIT.

integrated modeling↗

Optimizing fluvial flood mitigation strategies: A multi-objective approach for cost-effective and socially-aware infrastructure feasibility analysis

Effective levee planning must balance capital cost, risk reduction, and community priorities. These objectives are rarely optimized together. This study presents a feasibility phase, simulationin-the-loop framework that couples terrain-based flood modeling with a socially aware multiobjective optimizer. Flood risk is measured as Expected Annual Exposed Population (EAEP), obtained by integrating exposure over Annual Exceedance Probability (AEP) nodes, mirroring the Hydrologic Engineering Center's Flood Damage Reduction Analysis (HEC-FDA) expected-annual formulation but with people rather than dollars. Exposure per scenario is computed by overlaying binary inundation masks with a population surface at the tract level. Distributional fairness is encoded through a Group Benefit Share (GBS) constraint that requires high-SVI tracts to receive at least a baseline share of annualized benefits. Capital cost is represented by a height-dependent unit-cost model suitable for screening. This study addresses the two-objective problem, minimize cost and expected annual exposure subject to the GBS constraint, using Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and leveraging Pareto front for feasibility phase decision making. Implemented with terrain-based flood modeling, GeoFlood, for rapid scenario evaluation, the framework is demonstrated in Southeast Texas. The results reveal clear trade-offs among cost, risk, and social benefits and identify non-dominated levee height configurations that satisfy the benefit-share floor. The contributions are a scalable decision support method that operationalizes expected annual population-based risk, embeds enforceable benefit-sharing guarantees, and uses lightweight simulation to explore large design spaces before higher fidelity design stages.

Flood mitigation↗

BeyondPlanck: VII. Bayesian estimation of gain and absolute calibration for cosmic microwave background experiments

We present a Bayesian calibration algorithm for cosmic microwave background (CMB) observations as implemented within the global end-to-end BEYONDPLANCK framework and applied to the Planck Low Frequency Instrument (LFI) data. Following the most recent Planck analysis, we decomposed the full time-dependent gain into a sum of three nearly orthogonal components: one absolute calibration term, common to all detectors, one time-independent term that can vary between detectors, and one time-dependent component that was allowed to vary between one-hour pointing periods. Each term was then sampled conditionally on all other parameters in the global signal model through Gibbs sampling. The absolute calibration is sampled using only the orbital dipole as a reference source, while the two relative gain components were sampled using the full sky signal, including the orbital and Solar CMB dipoles, CMB fluctuations, and foreground contributions. We discuss various aspects of the data that influence gain estimation, including the dipole-polarization quadrupole degeneracy and processing masks. Comparing our solution to previous pipelines, we find good agreement in general, with relative deviations of -0.67% (-0.84%) for 30 GHz, 0.12% (-0.04%) for 44 GHz and -0.03% (-0.64%) for 70 GHz, compared to Planck PR4 and Planck 2018, respectively. We note that the BEYONDPLANCK calibration was performed globally, which results in better inter-frequency consistency than previous estimates. Additionally, WMAP observations were used actively in the BEYONDPLANCK analysis, which both breaks internal degeneracies in the Planck data set and results in an overall better agreement with WMAP. Finally, we used a Wiener filtering approach to smoothing the gain estimates. We show that this method avoids artifacts in the correlated noise maps as a result of oversmoothing the gain solution, which is difficult to avoid with methods like boxcar smoothing, as Wiener filtering by construction maintains a balance between data fidelity and prior knowledge. Although our presentation and algorithm are currently oriented toward LFI processing, the general procedure is fully generalizable to other experiments, as long as the Solar dipole signal is available to be used for calibration.

79 ASTRONOMY AND ASTROPHYSICS↗

Algorithm That Synthesizes Other Algorithms for Hashing

An algorithm that includes a collection of several subalgorithms has been devised as a means of synthesizing still other algorithms (which could include computer code) that utilize hashing to determine whether an element (typically, a number or other datum) is a member of a set (typically, a list of numbers). Each subalgorithm synthesizes an algorithm (e.g., a block of code) that maps a static set of key hashes to a somewhat linear monotonically increasing sequence of integers. The goal in formulating this mapping is to cause the length of the sequence thus generated to be as close as practicable to the original length of the set and thus to minimize gaps between the elements. The advantage of the approach embodied in this algorithm is that it completely avoids the traditional approach of hash-key look-ups that involve either secondary hash generation and look-up or further searching of a hash table for a desired key in the event of collisions. This algorithm guarantees that it will never be necessary to perform a search or to generate a secondary key in order to determine whether an element is a member of a set. This algorithm further guarantees that any algorithm that it synthesizes can be executed in constant time. To enforce these guarantees, the subalgorithms are formulated to employ a set of techniques, each of which works very effectively covering a certain class of hash-key values. These subalgorithms are of two types, summarized as follows: Given a list of numbers, try to find one or more solutions in which, if each number is shifted to the right by a constant number of bits and then masked with a rotating mask that isolates a set of bits, a unique number is thereby generated. In a variant of the foregoing procedure, omit the masking. Try various combinations of shifting, masking, and/or offsets until the solutions are found. From the set of solutions, select the one that provides the greatest compression for the representation and is executable in the minimum amount of time. Given a list of numbers, try to find one or more solutions in which, if each number is compressed by use of the modulo function by some value, then a unique value is generated.

James, Mark↗

Detecting Low Surface Brightness Galaxies with Mask R-CNN

Low surface brightness galaxies (LSBGs), galaxies that are fainter than the dark night sky, are famously difficult to detect. However, studies of these galaxies are essential to improve our understanding of the formation and evolution of low-mass galaxies. In this work, we train a deep learning model using the Mask R-CNN framework on a set of simulated LSBGs inserted into images from the Dark Energy Survey (DES) Data Release 2 (DR2). This deep learning model is combined with several conventional image pre-processing steps to develop a pipeline for the detection of LSBGs. We apply this pipeline to the full DES DR2 coadd image dataset, and preliminary results show the detection of 22 large, high-quality LSBG candidates that went undetected by conventional algorithms. Furthermore, we find that Galactic cirrus represents the largest contaminant in our resulting candidate list.

Levy, Caleb↗

Machine Learning Algorithms for Aerosol and Cloud Detection Using CATS on the ISS

Clouds and aerosols are one of the largest uncertainties in understanding and forecasting the Earth’s changing climate system. The type and height of aerosols are important factors in determining the top-of-atmosphere (TOA) radiation budget, either direct reflection of solar radiation back to space and/or absorption of solar radiation. In addition to their impact on the Earth’s climate system, aerosols near the surface from wildfires, man-made pollution events, and dust storms are hazardous to human health. The phase and height of clouds also play a critical role in determining the role of clouds in the Earth’s climate system. Cirrus clouds in the upper troposphere can induce a significant daytime TOA warming effect, while liquid water clouds near the surface cause a large corresponding cooling effect. Lidar measurements provide accurate vertically resolved information about clouds and aerosols, including complex multi-layer scenes where passive sensors are challenged and at night, when passive sensors are unable to measure cloud and aerosol properties. The Cloud-Aerosol Transport System (CATS) is a lidar instrument that operated for 33 months on the International Space Station (ISS) at the 1064 nm wavelength to measure attenuated total backscatter and depolarization ratio. These fundamental measurements are used to derive “vertical feature mask” cloud and aerosol products, including layer top/base heights, layer geometrical thickness, aerosol type, and cloud phase. While space-based lidar systems like CATS provide cloud and aerosol vertical distributions that improve our understanding of the climate system, averaging of the daytime data from these sensors is required, at the expense of spatial resolution, to improve the daytime signal-to noise (SNR) and thus atmospheric layer detection. This presentation shows results from machine learning (ML) techniques that, when applied to CATS data: 1. improve the 1064 nm SNR 2. enable detection of atmospheric features during daytime with a horizontal resolution of 350 m or 5 km (compared to the 60 km required for standard CATS data products) 3. increase the number of atmospheric layers detected in the CATS data. A Convolutional Neural Network (CNN) trained using CATS standard data products also demonstrated the potential for improved cloud-aerosol discrimination, cloud phase, and aerosol typing compared to the operational CATS algorithms for cloud edges and complex near-surface scenes during daytime. The ML tools described in this paper can facilitate the development of smaller, low-cost lidar systems in the future and enable real-time accessibility of lidar data products from future lidar systems for monitoring and forecasting of hazardous events.

John Yorks↗

The classification of the Arctic Sea ice types and the determination of surface temperature using advanced very high resolution radiometer data

The accurate quantification of new ice and open water areas and surface temperatures within the sea ice packs is a key to the realistic parameterization of heat, moisture, and turbulence fluxes between ocean and atmosphere in the polar regions. Multispectral NOAA advanced very high resolution radiometer/2 (AVHRR/2) satellite images are analyzed to evaluate how effectively the data can be used to characterize sea ice in the Bering and Greenland seas, both in terms of surface type and physical temperature. The basis of the classification algorithm, which is developed using a late wintertime Bering Sea ice cover data, is that frequency distributions of 10.8- micrometers radiances provide four distinct peaks, represeting open water, new ice, young ice, and thick ice with a snow cover. The results are found to be spatially and temporally consistent. Possible sources of ambiguity, especially associated with wider temporal and spatial application of the technique, are discussed. An ice surface temperature algorithm is developed for the same study area by regressing thermal infrared data from 10.8- and 12.0- micrometers channels against station air temperatures, which are assumed to approximate the skin temperatures of adjacent snow and ice. The standard deviations of the results when compared with in situ data are about 0.5 K over leads and polynyas to about 0.5-1.5 K over thick ice. This study is based upon a set of in situ data limited in scope and coverage. Cloud masks are applied using a thresholding technique that utilizes 3.74- and 10.8- micrometers channel data. The temperature maps produced show coherence with surface features like new ice and leads, and consistency with corresponding surface type maps. Further studies are needed to better understand the effects of both the spatial and temporal variability in emissivity, aerosol and precipitable atmospheric ice particle distribution, and atmospheric temperature inversions.

Massom, Robert↗

High–throughput measurement of plant fitness traits with an object detection method using Faster R–CNN

Revealing the contributions of genes to plant phenotype is frequently challenging because loss-of-function effects may be subtle or masked by varying degrees of genetic redundancy. Such effects can potentially be detected by measuring plant fitness, which reflects the cumulative effects of genetic changes over the lifetime of a plant. However, fitness is challenging to measure accurately, particularly in species with high fecundity and relatively small propagule sizes such as Arabidopsis thaliana. An image segmentation-based method using the software ImageJ and an object detection-based method using the Faster Region-based Convolutional Neural Network (R-CNN) algorithm were used for measuring two Arabidopsis fitness traits: seed and fruit counts. The segmentation-based method was error-prone (correlation between true and predicted seed counts, r 2 = 0.849) because seeds touching each other were undercounted. By contrast, the object detection-based algorithm yielded near perfect seed counts (r 2 = 0.9996) and highly accurate fruit counts (r 2 = 0.980). Comparing seed counts for wild-type and 12 mutant lines revealed fitness effects for three genes; fruit counts revealed the same effects for two genes. Our study provides analysis pipelines and models to facilitate the investigation of Arabidopsis fitness traits and demonstrates the importance of examining fitness traits when studying gene functions.

59 BASIC BIOLOGICAL SCIENCES↗

DESIVAST: Catalogs of Low-redshift Voids Using Data from the DESI Data Release 1 Bright Galaxy Survey

We present three separate void catalogs created using a volume-limited sample of the DESI Data Release 1 Bright Galaxy Survey. We use the algorithms VoidFinder and V 2 to construct void catalogs out to a redshift of z = 0.24. Excluding voids affected by the boundaries of the survey, we obtain 1489 voids with VoidFinder, 389 with V 2 using REVOLVER pruning, and 297 with V 2 using VIDE pruning. Comparing our catalogs with overlapping Sloan Digital Sky Survey void catalogs, we find generally consistent void properties but significant differences in the void volume overlap, which we attribute to differences in the galaxy selection and survey masks. These catalogs are suitable for studying the variation in galaxy properties with cosmic environment and for cosmological studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Integration of Satellite-Derived Cloud Phase, Cloud Top Height, and Liquid Water Path into an Operational Aircraft Icing Nowcasting System

Operational products used by the U.S. Federal Aviation Administration to alert pilots of hazardous icing provide nowcast and short-term forecast estimates of the potential for the presence of supercooled liquid water and supercooled large droplets. The Current Icing Product (CIP) system employs basic satellite-derived information, including a cloud mask and cloud top temperature estimates, together with multiple other data sources to produce a gridded, three-dimensional, hourly depiction of icing probability and severity. Advanced satellite-derived cloud products developed at the NASA Langley Research Center (LaRC) provide a more detailed description of cloud properties (primarily at cloud top) compared to the basic satellite-derived information used currently in CIP. Cloud hydrometeor phase, liquid water path, cloud effective temperature, and cloud top height as estimated by the LaRC algorithms are into the CIP fuzzy logic scheme and a confidence value is determined. Examples of CIP products before and after the integration of the LaRC satellite-derived products will be presented at the conference.

Haggerty, Julie↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗