Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for scientific computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto↗

Computational Analysis of Coupled Geoscience Processes in Fractured and Deformable Media

Prediction of flow, transport, and deformation in fractured and porous media is critical to improving our scientific understanding of coupled thermal-hydrological-mechanical processes related to subsurface energy storage and recovery, nonproliferation, and nuclear waste storage. Especially, earth rock response to changes in pressure and stress has remained a critically challenging task. In this work, we advance computational capabilities for coupled processes in fractured and porous media using Sandia Sierra Multiphysics software through verification and validation problems such as poro-elasticity, elasto-plasticity and thermo-poroelasticity. We apply Sierra software for geologic carbon storage, fluid injection/extraction, and enhanced geothermal systems. We also significantly improve machine learning approaches through latent space and self-supervised learning. Additionally, we develop new experimental technique for evaluating dynamics of compacted soils at an intermediate scale. Overall, this project will enable us to systematically measure and control the earth system response to changes in stress and pressure due to subsurface energy activities.

58 GEOSCIENCES↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗

Scientific machine learning for modeling and simulating complex fluids

The formulation of rheological constitutive equations—models that relate internal stresses and deformations in complex fluids—is a critical step in the engineering of systems involving soft materials. While data-driven models provide accessible alternatives to expensive first-principles models and less accurate empirical models in many engineering disciplines, the development of similar models for complex fluids has lagged. The diversity of techniques for characterizing non-Newtonian fluid dynamics creates a challenge for classical machine learning approaches, which require uniformly structured training data. Consequently, early machine-learning based constitutive equations have not been portable between different deformation protocols or mechanical observables. Here, we present a data-driven framework that resolves such issues, allowing rheologists to construct learnable models that incorporate essential physical information, while remaining agnostic to details regarding particular experimental protocols or flow kinematics. These scientific machine learning models incorporate a universal approximator within a materially objective tensorial constitutive framework. By construction, these models respect physical constraints, such as frame-invariance and tensor symmetry, required by continuum mechanics. We demonstrate that this framework facilitates the rapid discovery of accurate constitutive equations from limited data and that the learned models may be used to describe more kinematically complex flows. This inherent flexibility admits the application of these “digital fluid twins” to a range of material systems and engineering problems. We illustrate this flexibility by deploying a trained model within a multidimensional computational fluid dynamics simulation—a task that is not achievable using any previously developed data-driven rheological equation of state.

Science & Technology - Other Topics↗

Artificial Intelligence-Enhanced, Multi-Level, Modular System Design

As Moore’s Law and Dennard Scaling come to an end, it is becoming increasingly important to develop non-von Neumann computing architectures that can perform low-power computing in the domains of scientific computing, artificial intelligence, embedded systems, and edge computing. Next-generation computing technologies, such as neuromorphic computing and quantum computing, have the potential to revolutionize computing. However, in order to make progress in these fields, it is necessary to fundamentally change the current computing paradigm by codesigning systems across all system level, from materials to software. Because skilled labor is limited in the field of next-generation computing, we are developing artificial intelligence-enhanced tools to automate the codesign and co-discovery of next-generation computers. Here, we develop a method called Modular and Multi-level MAchine Learning (MAMMAL) which is able to perform analog codesign and co-discovery across multiple system levels, spanning devices to circuits. We prototype MAMMAL by using it to design simple passive analog low-pass filters. We also explore methods to incorporate uncertainty quantification into MAMMAL and to accelerate MAMMAL by using emerging technologies, such as crossbar arrays. Ultimately, we believe that MAMMAL will enable rapid progress in developing next-generation computers by automating the codesign and co-discovery of electronic systems.

97 MATHEMATICS AND COMPUTING↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Global field reconstruction from sparse sensors with Veronoi tessellation-assisted deep learning

Achieving accurate and robust global situational awareness of a complex time-evolving field from a limited number of sensors has been a longstanding challenge. This reconstruction problem is especially difficult when sensors are sparsely positioned in a seemingly random or unorganized manner, which is often encountered in a range of scientific and engineering problems. Moreover, these sensors can be in motion and can become online or offline over time. The key leverage in addressing this scientific issue is the wealth of data accumulated from the sensors. As a solution to this problem, we propose a data-driven spatial field recovery technique founded on a structured grid-based deep-learning approach for arbitrary positioned sensors of any numbers. It should be noted that the naïve use of machine learning becomes prohibitively expensive for global field reconstruction and is furthermore not adaptable to an arbitrary number of sensors. In the present work, we consider the use of Voronoi tessellation to obtain a structured-grid representation from sensor locations enabling the computationally tractable use of convolutional neural networks. One of the central features of the present method is its compatibility with deep-learning based super-resolution reconstruction techniques for structured sensor data that are established for image processing. The proposed reconstruction technique is demonstrated for unsteady wake flow, geophysical data, and three-dimensional turbulence. The current framework is able to handle an arbitrary number of moving sensors, and thereby overcomes a major limitation with existing reconstruction methods. The presented technique opens a new pathway towards the practical use of neural networks for real-time global field estimation.

Fukami, Kai↗

Advanced Modeling of Beam Physics and Performance Optimization for Nuclear Physics Colliders

High energy colliders provide a critical tool in nuclear physics study by probing the fundamental structure and dynamics of matter. To maximize the potential of scientific discovery in nuclear physics study, it is important to optimize the parameters of these colliders to attain the best performance. The performance of a collider is typically measured by its integrated luminosity of colliding beams since the probability of a new event is proportional to the integrated luminosity. However, the achievable luminosity is limited by the electromagnetic interactions (beam-beam effects) of two colliding beams at higher energy, and the interplay between the space-charge effects and the beam-beam effects at lower energy. To achieve the best performance of a collider means to attain the highest luminosity of the collider with optimized collider parameters. Optimizing the collider’s machine parameters is both computationally and experimentally expensive. A fast and robust computational framework including beam-beam and space-charge effects will be critical to attaining the best performance of the collider. In this project, we will study the beam dynamics challenges, specifically the interplay of the space-charge and the beam-beam effects, and the machine tuning models for maximizing the performance of RHIC experiments. We will develop an advanced modeling framework based on first-principles physical simulations, lattice models and the state-of-the-art machine learning methods and apply this framework to performance improvement of the RHIC in operation. We will build data manipulation packages to connect the simulation data and the experimental data with the framework, develop a self-consistent hybrid model of space-charge and beam-beam effects, study underlying physics mechanisms, build surrogate models using the labeled data, integrate the models into the advanced modeling framework, and apply the framework to RHIC luminosity (STAR and sPHENIX) optimization. The success of this project would substantially improve the performance of existing and future colliders and increase the opportunity for scientific discovery.

43 PARTICLE ACCELERATORS↗

Convolutional Neural Network–Aided Temperature Field Reconstruction: An Innovative Method for Advanced Reactor Monitoring

In this study, the capabilities of a physics-informed convolutional neural network (CNN) for reconstructing the temperature field from a limited set of measurements taken at the boundaries of internal flows are demonstrated. Such an approach enables the development of less invasive monitoring methods for real-time plant diagnostics. As a test case, a Molten Salt Fast Reactor (MSFR) design was selected. This circulating fuel reactor has received interest from both scientific and industrial communities due to its intrinsic safety and sustainability. Molten salt flows in such reactors, however, can present highly localized temperature peaks that can induce significant thermal stresses onto the vessel walls. At these local maxima, the salt temperature may exceed a thousand kelvins, which makes a direct measurement challenging or even unfeasible. The proposed CNN algorithm allows one to detect indirectly such discontinuities through an accurate, albeit indirect, temperature measurement method during reactor operation. The datasets employed to train and test the machine learning models in the present work were generated with Nek5000, a computational fluid dynamics (CFD) code developed at Argonne National Laboratory. The CNN algorithm is trained with CFD results that span a set of MSFR operational power and flow ranges. Here, to demonstrate the efficacy of the algorithm, predictions are made for test cases contained within the training range but for which the CFD data were not used when training. Results demonstrate that the proposed technique properly characterizes temperature peaks and distributions within the domain for a broad range of scenarios.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Multisensor Agile Adaptive Sampling of Convective Storms Driven by Real-time Analytics

Convective storms vertically transport water vapor and condensate from Earth’s surface to the upper troposphere. Life on Earth is fundamentally linked to this transport which determines the hydrological cycle, and the intensity of severe weather responsible for the destruction of life and property. Despite advances in high-resolution modeling and better observational capabilities, the scientific community continues to be confronted with knowledge gaps about convective storms that limit our predictive capabilities. The ongoing developments in the high-resolution Energy Exascale Earth System Model (E3SM), large eddy simulations, and AI-based analytics to evaluate uncertainties are expected to provide a comprehensive framework for new scientific discovery. The model-experiment (MODEX) approach suggests that the aforementioned advancements in model development and AI-based inference techniques should be complemented by similar advancements in the experimental (observational) side so that the former does not outstrip the ability of the latter to provide meaningful constraints. What are the recent advancements in observations that will provide the necessary leap forward in improving our predictive capabilities? To address this question, we propose a new experimental paradigm called Multisensor Agile Adaptive Sampling (MAAS) that capitalizes on advancements in communications (5G), computational resources (edge/fog computing), sensor capabilities, and machine learning (ML) and AI techniques (Kollias et al., 2020). The MAAS framework allows for the collection of higher spatiotemporal resolution and quality observations of convective storms than is traditionally possible. The MAAS framework is scalable and applicable to atmospheric observatories such as those operated by the Department of Energy (DoE) Atmospheric Radiation Measurement (ARM) facility.

54 ENVIRONMENTAL SCIENCES↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

Investigating the Future of Scientific Data Search [Slides]

Searching for usable, actionable, data in a trustworthy manner is a challenge across scientific communities. Artificial Intelligence (AI) and Machine Learning (ML) techniques may be leveraged to increase the utility of scientific data by: Demystify unstructured data to aid curation & sharing Surfacing hard to find datasets. User Experience (UX) Research can help uncover scientists needs & challenges finding data and using AI/ML enabled tools.

97 MATHEMATICS AND COMPUTING↗

Zentropy Theory for Transformative Functionalities of Magnetic and Superconducting Materials

The proposed research developed the zentropy theory through applications to complex magnetic materials and superconductors under the hypothesis that the emergent properties of complex magnetic materials and superconductors can be predicted by statistical mechanics of ergodic microstates with their partition functions computed from DFT-predicted free energies. The key objective is to develop approaches to systematically determine the types and number of microstates and the supercell size in DFT-based calculations through convergency of macroscopic functionalities, with the incorporation of our mixed-space approach accounting for the interactions between periodic supercells. In addition to use scientific intuitions to guide the design of important microstates, the key innovation of the proposed research is to integrate the domain knowledge and the material-property-descriptor database (MPDD) with 4 million microstates, which is supported by our deep neural network machine learning models (SIPFENN: structure-informed prediction of formation energy using neural networks) and integrated with our high throughput DFT Tool Kit (DFTTK). For complex magnetic materials, one of the objectives is to develop approaches to calculate short-range ordering from the statistical distribution of each microstate. For superconductors, the divergency of quasiparticle effective mass at a quantum critical point will be investigated, and the superconducting and non-superconducting microstates will be delineated through analysis of electronic band structure, density of states, charge density, and Fermi surface.

36 MATERIALS SCIENCE↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗