Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

The Virtual Blast Furnace - An Integrated High Performance Computing Modeling, Simulation, and Visualization Capability for Steel Manufacturing (Final Report)

Many manufacturing industries require substantial capital and utilize energy intensive processes that involve complex phenomena. One example of such an industry is the steel industry, which is the fourth largest energy consuming industry in the U.S. By harnessing the power of High-Performance Computing (HPC) to enhance current simulation and visualization methods in the steel industry, it should be possible to increase resolution and/or decrease time of these methods by a factor of 1000. In this way, information can be obtained in a time frame that is useful for making business and engineering decisions, optimizing manufacturing processes and, ultimately, improving the completeness of U.S. industries. For example, if coke usage in blast furnaces were optimized such that the average coke rate was reduced from 797 lb/net tonne of hot metal (NTHM) to 604 lb/NTHM, costs could be reduced by $894 million/year. Additionally, members of the steel industry need the flexibility to efficiently operate blast furnaces at a range of production rates in order to meet fluctuating market demands. Large scale parameter studies can be utilized to discover workable operating parameters at a range of production rates.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A critical examination of compound stability predictions from machine-learned formation energies

Machine learning has emerged as a novel tool for the efficient prediction of material properties, and claims have been made that machine-learned models for the formation energy of compounds can approach the accuracy of Density Functional Theory (DFT). The models tested in this work include five recently published compositional models, a baseline model using stoichiometry alone, and a structural model. By testing seven machine learning models for formation energy on stability predictions using the Materials Project database of DFT calculations for 85,014 unique chemical compositions, we show that while formation energies can indeed be predicted well, all compositional models perform poorly on predicting the stability of compounds, making them considerably less useful than DFT for the discovery and design of new solids. Most critically, in sparse chemical spaces where few stoichiometries have stable compounds, only the structural model is capable of efficiently detecting which materials are stable. The nonincremental improvement of structural models compared with compositional models is noteworthy and encourages the use of structural models for materials discovery, with the constraint that for any new composition, the ground-state structure is not known a priori. This work demonstrates that accurate predictions of formation energy do not imply accurate predictions of stability, emphasizing the importance of assessing model performance on stability predictions, for which we provide a set of publicly available tests.

36 MATERIALS SCIENCE↗

Nonlinear modeling and performance analysis of cracked beam microgyroscopes

Cracks constitute a common structural defect in microelectromechanical systems that may arise during manufacturing, mechanical fatigue, or shock loading. Areas weakened by cracks increase the flexibility of the microstructure. Furthermore, depending on the crack severity, the increased flexibility can cause dramatic changes in the static and dynamic behaviors of electrically actuated MEMS. In this study, a numerical investigation of the depth and location of multiple edge cracks on the performance of beam microgyroscopes is conducted. Applying the Griffith strain energy release theorem for an edge crack, the vibrational characteristics of microgyroscopes with multiple cracks including the static deflection, natural frequencies, and mode shapes are studied analytically. Then, the differential quadrature method is employed to model the nonlinear dynamic behaviors of the cracked gyroscope system. Cracks initiated along the driving or sensing directions of the microgyroscope are found to have differing effects on its nonlinear dynamic response. This numerical study reveals that the onset of severe single cracks or a series of multiple cracks can significantly affect the sensitivity of the microgyroscope to base rotations and therefore leads to a degradation in the performance of the damaged MEMS sensor. This study also shows that a potential design approach for broadband microgyroscope systems manufactured with cracks can be proposed.

42 ENGINEERING↗

A Surface Radiation Balance Dataset from Siple Dome in West Antarctica for Atmospheric and Climate Model Evaluation

Abstract A field campaign at Siple Dome in West Antarctica during the austral summer 2019/20 offers an opportunity to evaluate climate model performance, particularly cloud microphysical simulation. Over Antarctic ice sheets and ice shelves, clouds are a major regulator of the surface energy balance, and in the warm season their presence occasionally induces surface melt that can gradually weaken an ice shelf structure. This dataset from Siple Dome, obtained using transportable and solar-powered equipment, includes surface energy balance measurements, meteorology, and cloud remote sensing. To demonstrate how these data can be used to evaluate model performance, comparisons are made with meteorological reanalysis known to give generally good performance over Antarctica (ERA5). Surface albedo measurements show expected variability with observed cloud amount, and can be used to evaluate a model’s snowpack parameterization. One case study discussed involves a squall with northerly winds, during which ERA5 fails to produce cloud cover throughout one of the days. A second case study illustrates how shortwave spectroradiometer measurements that encompass the 1.6- μ m atmospheric window reveal cloud phase transitions associated with cloud life cycle. Here, continuously precipitating mixed-phase clouds become mainly liquid water clouds from local morning through the afternoon, not reproduced by ERA5. We challenge researchers to run their various regional or global models in a manner that has the large-scale meteorology follow the conditions of this field campaign, compare cloud and radiation simulations with this Siple Dome dataset, and potentially investigate why cloud microphysical simulations or other model components might produce discrepancies with these observations. significance statement Antarctica is a critical region for understanding climate change and sea level rise, as the great ice sheets and the ice shelves are subject to increasing risk as global climate warms. Climate models have difficulties over Antarctica, particularly with simulation of cloud properties that regulate snow surface melting or refreezing. Atmospheric and climate-related field work has significant challenges in the Antarctic, due to the small number of research stations that can support state-of-the-art equipment. Here we present new data from a suite of transportable and solar-powered instruments that can be deployed to remote Antarctic sites, including regions where ice shelves are most at risk, and we demonstrate how key components of climate model simulations can be evaluated against these data.

54 ENVIRONMENTAL SCIENCES↗

Impact of Color Space and Color Resolution on Vehicle Recognition Models

In this study, we analyze both linear and nonlinear color mappings by training on versions of a curated dataset collected in a controlled campus environment. We experiment with color space and color resolution to assess model performance in vehicle recognition tasks. Color encodings can be designed in principle to highlight certain vehicle characteristics or compensate for lighting differences when assessing potential matches to previously encountered objects. The dataset used in this work includes imagery gathered under diverse environmental conditions, including daytime and nighttime lighting. Experimental results inform expectations for possible improvements with automatic color space selection through feature learning. Moreover, we find there is only a gradual decrease in model performance with degraded color resolution, which suggests the need for simplified data collection and processing. By focusing on the most critical features, we could see improved model generalization and robustness, as the model becomes less prone to overfitting to noise or irrelevant details in the data. Such a reduction in resolution will lower computational complexity, leading to quicker training and inference times.

47 OTHER INSTRUMENTATION↗

Automated prediction of lattice parameters from X-ray powder diffraction patterns

A key step in the analysis of powder X-ray diffraction (PXRD) data is the accurate determination of unit-cell lattice parameters. This step often requires significant human intervention and is a bottleneck that hinders efforts towards automated analysis. This work develops a series of one-dimensional convolutional neural networks (1D-CNNs) trained to provide lattice parameter estimates for each crystal system. A mean absolute percentage error of approximately 10% is achieved for each crystal system, which corresponds to a 100- to 1000-fold reduction in lattice parameter search space volume. The models learn from nearly one million crystal structures contained within the Inorganic Crystal Structure Database and the Cambridge Structural Database and, due to the nature of these two complimentary databases, the models generalize well across chemistries. A key component of this work is a systematic analysis of the effect of different realistic experimental non-idealities on model performance. It is found that the addition of impurity phases, baseline noise and peak broadening present the greatest challenges to learning, while zero-offset error and random intensity modulations have little effect. However, appropriate data modification schemes can be used to bolster model performance and yield reasonable predictions, even for data which simulate realistic experimental non-idealities. In order to obtain accurate results, a new approach is introduced which uses the initial machine learning estimates with existing iterative whole-pattern refinement schemes to tackle automated unit-cell solution.

42 ENGINEERING↗

Search for HH → bbτ⁺τ⁻ Using Run 3 Scouting Data Analyze b-tagging and tau-tagging Performance with Unified Particle Transformer

B-tagging and tau-tagging performances play an important role in the search for the rare event HH → bbτ⁺τ⁻. A transformer-based neural network, Unified Particle Transformer, is applied for both tagging tasks, and Run 3 proton–proton collision scouting data at center-of-mass energy of 13.6 TeV is used. The scouting data stream accepts events at a much higher rate compared to traditional triggers, but stores only the objects reconstructed in the trigger, no low-level detector information. Therefore, existing taggers trained for the offline event reconstruction cannot be used. Analysis of the SoftMax plots, ROC/AUC curves, confusion matrix, accuracy and losses are used to evaluate model performance. Specifically, the tagging efficiency of the signal and misidentification probability across multiple background processes are compared for varying working points. Different training samples with distinct distributions of jet flavors are utilized and related model performances are analyzed. Interpretability methods, such as Integrated Gradients, may further be applied to study the input features’ influence on the model’s decisions, providing insights into potential improvements.

Chen, Blair [Purdue U., West Lafayette; Fermilab]↗

Multimodel validation of single wakes in neutral and stratified atmospheric conditions

Previous research has revealed the need for a validation study that considers several wake quantities and code types so that decisions on the trade-off between accuracy and computational cost can be well informed and appropriate to the intended application. In addition to guiding code choice and setup, rigorous model validation exercises are needed to identify weaknesses and strengths of specific models and guide future improvements. Here, we consider 13 approaches to simulating wakes observed with a nacelle-mounted lidar at the Scaled Wind Technology Facility (SWiFT) under varying atmospheric conditions. We find that some of the main challenges in wind turbine wake modeling are related to simulating the inflow. In the neutral benchmark, model performance tracked as expected with model fidelity, with large-eddy simulations performing the best. In the more challenging stable case, steady-state Reynolds-averaged Navier–Stokes simulations were found to outperform other model alternatives because they provide the ability to more easily prescribe noncanonical inflows and their low cost allows for simulations to be repeated as needed. Dynamic measurements were only available for the unstable benchmark at a single downstream distance. These dynamic analyses revealed that differences in the performance of time-stepping models come largely from differences in wake meandering. This highlights the need for more validation exercises that take into account wake dynamics and are able to identify where these differences come from: mesh setup, inflow, turbulence models, or wake-meandering parameterizations. In addition to model validation findings, we summarize lessons learned and provide recommendations for future benchmark exercises.

17 WIND ENERGY↗

NETL Plastic Pipes Project (Final Report)

Plastic or composite pipelines have been the bane of the utility locating industry because they are neither conductive nor magnetic which are the properties traditionally used to locate buried utilities. Ground penetrating radar (GPR) is an effective geophysical tool for locating plastic/composite pipelines where resistive cover allows for adequate penetration of radar energy. However, GPR has limited applicability in areas where the soil cover is conductive due to significant clay and/or salt content. This study examines complementary near-surface geophysical methods that are potentially useful for locating buried plastic/composite pipelines, either singly or in combination. Specifically, this modeling study used computational numerical methods to forward model the response of GPR, resistivity, seismic, gravity gradiometry, and photoacoustic/thermoacoustic imaging methods to plastic/composite pipelines for various scenarios including: (1) pipe diameters ranging between 2 in. to 12 in.; (2) burial depths ranging between 3 ft. to 4 ft.; (3) various degrees in contrast in physical properties (i.e., electrical permittivity, elasticity, resistivity, density); and (4) various experimental acquisition choices (e.g., GPR radar and seismic source frequencies, electrode spacing). Numerical modeling performed herein reconfirmed that GPR is the preferred method for detecting/locating plastic pipelines. A caveat for GPR detection is that the material covering the plastic pipe (trench fill material and adjacent soil) must be sufficiently resistive to allow the two-way propagation to the required depth of investigation and back to the surface. GPR was the only method modeled in this study that can be used to directly detect plastic pipelines of 2-in.-diameter and larger when buried 3-ft-deep. GPR data processing and imaging also can determine pipe depth, pipe diameter, trench dimensions, and moisture conditions. Seismic modeling results suggest that direct detection of a 12-in.-diameter plastic pipe at 3-ft.-depth may be possible under favorable conditions; however, the associated signature would be weak (e.g., surface- to S-wave, backscattered surface-waves, and/or forward scattered surface-waves to S-wave). Direct pipe detection under field conditions with noise and strong lateral geologic heterogeneity is doubtful. Numerical modeling also suggests that plastic pipelines can be indirectly located by detecting the trench in which they are buried. GPR, direct current (DC) resistivity, and seismic methods have the potential to locate the pipeline trench if there is sufficient contrast between the trench-wall and trench-fill materials for the physical property being measured by each method (i.e., electrical permittivity for GPR; resistivity for DC resistivity; or density, compressional velocity, or shear velocity for seismic). Modeling also indicated that currently available (commercial) gravity gradiometers would be unable to directly detect/locate plastic pipelines ≤ 8-in.-diameter when buried 3-ft.-deep given the typical instrument noise floor for field surveying as well as the expected density variations due to geologic heterogeneity. The numerical modeling performed in this project did not identify a universal geophysical technology that can locate buried plastic pipelines in all parts of the United States (although GPR is suggested for all areas with resistive cover). However, the project results suggest that a towed land streamer simultaneously acquiring multiple geophysical data types including multi-offset GPR, multi-channel DC resistivity, seismic geophone- and/or distributed acoustic sensing (DAS), and potentially photoacoustic/thermoacoustic data would be an appropriate platform for locating buried plastic pipeline. Moreover, the complementary multiphysics data acquired by a towed land streamer would permit the use of joint and/or cooperative inversion frameworks for a more rigorous and consistent data interpretation.

42 ENGINEERING↗

Modeling Subsurface Performance of a Geothermal Reservoir Using Machine Learning

Geothermal power plants typically show decreasing heat and power production rates over time. Mitigation strategies include optimizing the management of existing wells—increasing or decreasing the fluid flow rates across the wells—and drilling new wells at appropriate locations. The latter is expensive, time-consuming, and subject to many engineering constraints, but the former is a viable mechanism for periodic adjustment of the available fluid allocations. In this study, we describe a new approach combining reservoir modeling and machine learning to produce models that enable such a strategy. Our computational approach allows us, first, to translate sets of potential flow rates for the active wells into reservoir-wide estimates of produced energy, and second, to find optimal flow allocations among the studied sets. In our computational experiments, we utilize collections of simulations for a specific reservoir (which capture subsurface characterization and realize history matching) along with machine learning models that predict temperature and pressure timeseries for production wells. We evaluate this approach using an “open-source” reservoir we have constructed that captures many of the characteristics of Brady Hot Springs, a commercially operational geothermal field in Nevada, USA. Selected results from a reservoir model of Brady Hot Springs itself are presented to show successful application to an existing system. In both cases, energy predictions prove to be highly accurate: all observed prediction errors do not exceed 3.68% for temperatures and 4.75% for pressures. In a cumulative energy estimation, we observe prediction errors that are less than 4.04%. A typical reservoir simulation for Brady Hot Springs completes in approximately 4 h, whereas our machine learning models yield accurate 20-year predictions for temperatures, pressures, and produced energy in 0.9 s. This paper aims to demonstrate how the models and techniques from our study can be applied to achieve rapid exploration of controlled parameters and optimization of other geothermal reservoirs.

15 GEOTHERMAL ENERGY↗

Machine Learning Calibration of Groundwater Table Depth in ELM: Impact on Land Surface Hydrology and Land‐Atmosphere Fluxes

Accurate representation of groundwater table depth (GWTD) is crucial for simulating hydrological cycling in Earth system models (ESM). Nevertheless, there is a notable gap in the literature regarding the validation of GWTD simulations in ESMs and their subsequent impact on downstream hydrological components. This study explores the calibration of parameterization of global GWTD using machine learning within the Energy Exascale Earth System Model (E3SM) Land Model (ELM). Despite achieving significant gains in simulating GWTD through calibration, offline ELM simulations unexpectedly show that these improvements do not translate to substantial enhancements in model performance for other key hydrological variables, including soil moisture (SM), runoff, groundwater contribution to runoff or base flow index (BFI), and evapotranspiration and its partitioning. The performance in SM and runoff was even degraded in some regions, while BFI was mostly overestimated. Although there is significant improvement in GWTD within the critical range of 1–5 m, where groundwater traditionally influences land surface energy fluxes, these improvements occurred mostly in humid areas where the impact of GWTD on surface processes is minimal. Although the impacts of model calibration are generally small in offline ELM simulations, coupled land-atmosphere simulations exhibit much stronger responses to GWTD calibration, highlighting the role of land-atmosphere feedbacks in Earth system modeling. These findings underscore the need for integrated calibration strategies that simultaneously optimize multiple hydrological variables. However, if a single-variable approach is necessary, it is crucial to establish clear priorities for calibration, identifying the most critical variables that have the greatest impact on overall model performance.

Fang, Yilin [Pacific Northwest National Laboratory↗

Improving protein tertiary structure prediction by deep learning and distance prediction in CASP14

Abstract Substantial progresses in protein structure prediction have been made by utilizing deep‐learning and residue‐residue distance prediction since CASP13. Inspired by the advances, we improve our CASP14 MULTICOM protein structure prediction system by incorporating three new components: (a) a new deep learning‐based protein inter‐residue distance predictor to improve template‐free (ab initio) tertiary structure prediction, (b) an enhanced template‐based tertiary structure prediction method, and (c) distance‐based model quality assessment methods empowered by deep learning. In the 2020 CASP14 experiment, MULTICOM predictor was ranked seventh out of 146 predictors in tertiary structure prediction and ranked third out of 136 predictors in inter‐domain structure prediction. The results demonstrate that the template‐free modeling based on deep learning and residue‐residue distance prediction can predict the correct topology for almost all template‐based modeling targets and a majority of hard targets (template‐free targets or targets whose templates cannot be recognized), which is a significant improvement over the CASP13 MULTICOM predictor. Moreover, the template‐free modeling performs better than the template‐based modeling on not only hard targets but also the targets that have homologous templates. The performance of the template‐free modeling largely depends on the accuracy of distance prediction closely related to the quality of multiple sequence alignments. The structural model quality assessment works well on targets for which enough good models can be predicted, but it may perform poorly when only a few good models are predicted for a hard target and the distribution of model quality scores is highly skewed. MULTICOM is available at https://github.com/jianlin-cheng/MULTICOM_Human_CASP14/tree/CASP14_DeepRank3 and https://github.com/multicom-toolbox/multicom/tree/multicom_v2.0 .

59 BASIC BIOLOGICAL SCIENCES↗

High-Current Density Durability of Pt/C and PtCo/C Catalysts at Similar Particle Sizes in PEMFCs

The durability of carbon supported PtCo-alloy based nanoparticle catalysts play a key role in the longevity of proton-exchange membrane fuel cells (PEMFC) in electric vehicle applications. To improve its durability, it is important to understand and mitigate the various factors that cause PtCo-based cathode catalyst layers (CCL) to lose performance over time. These factors include i) electrochemical surface area (ECSA) loss, ii) specific activity loss, iii) H + /O 2 -transport changes and iv) Co 2+ contamination effects. We use a catalyst-specific accelerated stress test (AST) voltage cycling protocol to compare the durability of Pt and PtCo catalysts at similar average nanoparticle size and distribution. Our studies indicate that while Pt and PtCo nanoparticle catalysts suffer from similar magnitudes of electrochemical surface area (ECSA) losses, PtCo catalyst shows a significantly larger cell voltage loss at high current densities upon durability testing. The distinctive factor causing the large cell voltage loss of PtCo catalyst appears to be the secondary effects of the leached Co 2+ cations that contaminate the electrode ionomer. A 1D performance model has been used to quantify the cell voltage losses arising from various factors causing degradation of the membrane electrode assembly (MEA).

08 HYDROGEN↗

Updating and Evaluating Anthropogenic Emissions for NOAA’s Global Ensemble Forecast Systems for Aerosols (GEFS-Aerosols): Application of an SO 2 Bias-Scaling Method

We updated the anthropogenic emissions inventory in NOAA’s operational Global Ensemble Forecast for Aerosols (GEFS-Aerosols) to improve the model’s prediction of aerosol optical depth (AOD). We used a methodology to quickly update the pivotal global anthropogenic sulfur dioxide (SO 2 ) emissions using a speciated AOD bias-scaling method. The AOD bias-scaling method is based on the latest model predictions compared to NASA’s Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA2). The model bias was subsequently applied to the CEDS 2019 SO 2 emissions for adjustment. The monthly mean GEFS-Aerosols AOD predictions were evaluated against a suite of satellite observations (e.g., MISR, VIIRS, and MODIS), ground-based AERONET observations, and the International Cooperative for Aerosol Prediction (ICAP) ensemble results. The results show that transitioning from CEDS 2014 to CEDS 2019 emissions data led to a significant improvement in the operational GEFS-Aerosols model performance, and applying the bias-scaled SO 2 emissions could further improve global AOD distributions. The biases of the simulated AODs against the observed AODs varied with observation type and seasons by a factor of 3~13 and 2~10, respectively. The global AOD distributions showed that the differences in the simulations against ICAP, MISR, VIIRS, and MODIS were the largest in March–May (MAM) and the smallest in December–February (DJF). When evaluating against the ground-truth AERONET data, the bias-scaling methods improved the global seasonal correlation (r), Index of Agreement (IOA), and mean biases, except for the MAM season, when the negative regional biases were exacerbated compared to the positive regional biases. The effect of bias-scaling had the most beneficial impact on model performance in the regions dominated by anthropogenic emissions, such as East Asia. However, it showed less improvement in other areas impacted by the greater relative transport of natural emissions sources, such as India. The accuracies of the reference observation or assimilation data for the adjusted inputs and the model physics for outputs, and the selection of regions with less seasonal emissions of natural aerosols determine the success of the bias-scaling methods. A companion study on emission scaling of anthropogenic absorbing aerosols needs further improved aerosol prediction.

54 ENVIRONMENTAL SCIENCES↗

Comparison of Model Predictions and Performance Test Data for a Prototype Thermal Energy Storage Module

Although model predictions of thermal energy storage (TES) performance have been explored in previous investigations, relevant test data that enable experimental validation of performance models have been limited. This is particularly true for high-performance TES designs that facilitate fast input and extraction of energy. In this paper, we present a summary of experimental tests of a high-performance TES unit using lithium nitrate trihydrate phase change material as a storage medium. Performance data are presented for complete dual-mode cycles consisting of extraction (melting) followed by charging (freezing). These tests simulate the cyclic operation of a TES unit for asynchronous cooling in a variety of applications. Finally, the model analysis is found to agree reasonably well, within 10%, with the experimental data except for conditions very near the initiation of freezing, a consequence of subcooling that is required to initiate solidification.

25 ENERGY STORAGE↗

Isomeric effects on the reactivity of branched alkenes: An experimental and kinetic modeling study of methylbutenes

Here, a detailed experimental study of the low-to-intermediate temperature combustion of methylbutene isomers, i.e., branched C 5 alkenes, has been undertaken with multiple experimental facilities. Ignition delay times were measured at equivalence ratios 0.5–2.0, 685–1020 K and up to 45 bar condition from two rapid compression machines and showed slight deviation from an Arrhenius behavior for all three isomers, while their reactivity order differs as temperature changes. Sampled intermediates formed during the oxidation process of mixtures at 900–1150 K and 0.82 bar from a flow reactor and at 730 K and 20 bar from a rapid compression machine were analyzed using gas chromatography techniques. Trends in the formation and consumption of sampled intermediates were modeled using a kinetic model developed in this work for all three isomers. Rate of production and sensitivity analyses emphasize the role of double bond-specific reactions governing the global reactivity of these fuels. Additional studies of the addition reactions of HO 2 radicals to the double bond and to allylic radicals may improve the model performance.

2-Methyl-1-butene↗

Semisupervised Learning for Seismic Monitoring Applications

The impressive performance that deep neural networks demonstrate on a range of seismic monitoring tasks depends largely on the availability of event catalogs that have been manually curated over many years or decades. However, the quality, duration, and availability of seismic event catalogs vary significantly across the range of monitoring operations, regions, and objectives. Semisupervised learning (SSL) enables learning from both labeled and unlabeled data and provides a framework to leverage the abundance of unreviewed seismic data for training deep neural networks on a variety of target tasks. We apply two SSL algorithms (mean-teacher and virtual adversarial training) as well as a novel hybrid technique (exponential average adversarial training) to seismic event classification to examine how unlabeled data with SSL can enhance model performance. In general, we find that SSL can perform as well as supervised learning with fewer labels. We also observe in some scenarios that almost half of the benefits of SSL are the result of the meaningful regularization enforced through SSL techniques and may not be attributable to unlabeled data directly. Lastly, the benefits from unlabeled data scale with the difficulty of the predictive task when we evaluate the use of unlabeled data to characterize sources in new geographic regions. Finally, in geographic areas where supervised model performance is low, SSL significantly increases the accuracy of source-type classification using unlabeled data.

58 GEOSCIENCES↗

Analysis of GPU Data Access Patterns on Complex Geometries for the D3Q19 Lattice Boltzmann Algorithm

GPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When implemented with GPUs on complex domains, typically, geometric data is accessed indirectly and lattice data is accessed lexicographically. Although there are a variety of other options, no study has examined the relative efficacy between them. Here, we examine a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct is often better suited than the more common indirect addressing, providing increased computational speed and reducing memory consumption. For the layout, we find that the Collected Structure of Arrays (CSoA) and bundling layouts outperform the common Structure of Array layout; on V100 and P100 devices, CSoA consistently outperforms bundling, however the relationship is more complicated on K40 devices. When compared to state-of-the-art practices, our recommendations lead to speedups of 10–40 percent and reduce memory consumption up to 17 percent. Using performance modeling and computational experimentation, we determine the mechanisms behind the accelerations. We demonstrate that our results hold across multiple GPUs on two leadership class systems, and present the first near-optimal strong results for LBM with arterial geometries run on GPUs.

42 ENGINEERING↗