Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Optimization of Thermal Conductance at Interfaces Using Machine Learning Algorithms

We report optimization of thermal transport across the interface of two different materials is critical to micro-/nanoscale electronic, photonic, and phononic devices. Although several examples of compositional intermixing at the interfaces having a positive effect on interfacial thermal conductance (ITC) have been reported, an optimum arrangement has not yet been determined because of the large number of potential atomic configurations and the significant computational cost of evaluation. On the other hand, computation-driven materials design efforts are rising in popularity and importance. Yet, the scalability and transferability of machine learning models remain as challenges in creating a complete pipeline for the simulation and analysis of large molecular systems. In this work we present a scalable Bayesian optimization framework, which leverages dynamic spawning of jobs through the Message Passing Interface (MPI) to run multiple parallel molecular dynamics simulations within a parent MPI job to optimize heat transfer at the silicon and aluminum (Si/Al) interface. We found a maximum of 50% increase in the ITC when introducing a two-layer intermixed region that consists of a higher percentage of Si. Because of the random nature of the intermixing, the magnitude of increase in the ITC varies. We observed that both homogeneity/heterogeneity of the intermixing and the intrinsic stochastic nature of molecular dynamics simulations account for the variance in ITC.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗

Direction-optimizing Label Propagation and its Application for Community Detection

Label Propagation is a machine learning algorithm typically used for classification. It has also been found to be an effective method for detecting communities in networks. It has two attractive features as a community detection method: it has nearly linear runtime and it requires no \textit{a priori} community information. We propose a new Direction Optimizing Label Propagation Algorithm (DOLPA) that relies on the use of {\em frontiers} and alternates between label {\em push} and label {\em pull} operations to enhance the performance of LPA. Specifically, DOLPA has parameters for tuning the processing order of vertices in a graph. This reduces the number of edges visited and improves the quality of solution. We apply DOLPA to community detection and present the design and implementation of the algorithm as well as its shared-memory parallelization using OpenMP. Empirically, we evaluate our algorithm using synthetic graphs as well as real-world networks. Compared with the state-of-art \textit{Parallel Label Propagation} algorithm, we achieve at least two times the F-Score while reducing the runtime by 50\% for synthetic graphs with overlapping communities. We also compare DOLPA against the state-of-art parallel implementation of the Louvain method using the same graphs and show that DOLPA achieves about three times the F-Score at 10\% the runtime. On real-world graphs, we get a speedup of up to $10\times$ using 64 threads.

Liu, X↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

Reaction Mechanism Generator v3.0: Advances in Automatic Mechanism Generation

In chemical kinetics research, kinetic models containing hundreds of species and tens of thousands of elementary reactions are commonly used to understand and predict the behavior of reactive chemical systems. Reaction Mechanism Generator (RMG) is a software suite developed to automatically generate such models by incorporating and extrapolating from a database of known thermochemical and kinetic parameters. Here, we present the recent version 3 release of RMG and highlight improvements since the previously published description of RMG v1.0. Most notably, RMG can now generate heterogeneous catalysis models in addition to the previously available gas- and liquid-phase capabilities. For model analysis, new methods for local and global uncertainty analysis have been implemented to supplement first-order sensitivity analysis. The RMG database of thermochemical and kinetic parameters has been significantly expanded to cover more types of chemistry. The present release includes parallelization for faster model generation and a new molecule isomorphism approach to improve computational performance. RMG has also been updated to use Python 3, ensuring compatibility with the latest cheminformatics and machine learning packages. Overall, RMG v3.0 includes many changes which improve the accuracy of the generated chemical mechanisms and allow for exploration of a wider range of chemical systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Extreme-scale stochastic optimization and simulation via learning-enhanced decomposition and parallelization (Final Technical Report)

Stochastic optimization and simulation models ubiquitously arise in designing and operating complex service/engineering systems. They can be extreme in scale due to high-dimensional data and decisions, and can also involve decisions made sequentially in response to newly revealed data, both causing significant computational challenge. The objective of this research is to explore a unified framework that integrates machine learning with discrete optimization and risk-averse modeling, to improve the efficiency of decomposition paradigms for stochastic optimization and simulations at extreme scale. The models we consider represent a broad class of complex decision-making problems, where 0-1 or continuous decisions are made before and/or after knowing multiple sources of uncertainties that could be correlated. We will employ machine learning methods to dynamically decide and prioritize computational procedures, including cut generation, branching, and bounding of the optimal objective. Furthermore, the research will shed new lights on the traditional decomposition algorithms for extreme-scale computing. Deliverables of the research include new modeling and computational methods for advancing the state-of-the-art research in optimization and simulation, bringing many relevant risk-averse, data-driven optimization problems in practice within the range of tractability. Examples include distributed computing server scheduling and sensor deployment for monitoring critical infrastructures. Success in this effort will enable progress in solving multiple extreme-scale problems in the complex system design and operations arising from DoE missions in energy, environment, and national security.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Combining machine-learned and empirical force fields with the parareal algorithm: application to the diffusion of atomistic defects

We numerically investigate an adaptive version of the parareal algorithm in the context of molecular dynamics. This adaptive variant has been originally introduced in [1]. We focus here on test cases of physical interest where the dynamics of the system is modelled by the Langevin equation and is simulated using the molecular dynamics software LAMMPS. In this work, the parareal algorithm uses a family of machine-learning spectral neighbor analysis potentials (SNAP) as fine, reference, potentials and embedded-atom method potentials (EAM) as coarse potentials. We consider a self-interstitial atom in a tungsten lattice and compute the average residence time of the system in metastable states. Our numerical results demonstrate significant computational gains using the adaptive parareal algorithm in comparison to a sequential integration of the Langevin dynamics. We also identify a large regime of numerical parameters for which statistical accuracy is reached without being a consequence of trajectorial accuracy.

36 MATERIALS SCIENCE↗

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

Direction-optimizing Label Propagation Framework for Structure Detection in Graphs: Design, Implementation, and Experimental Analysis

Label Propagation is not only a well-known machine learning algorithm for classification but also an effective method for discovering communities and connected components in networks. We propose a new Direction-optimizing Label Propagation Algorithm (DOLPA) framework that enhances the performance of the standard Label Propagation Algorithm (LPA), increases its scalability, and extends its versatility and application scope. As a central feature, the DOLPA framework relies on the use of frontiers and alternates between label push and label pull operations to attain high performance. It is formulated in such a way that the same basic algorithm can be used for finding communities or connected components in graphs by only changing the objective function used. Additionally, DOLPA has parameters for tuning the processing order of vertices in a graph to reduce the number of edges visited and improve the quality of solution obtained. We present the design and implementation of the enhanced algorithm as well as our shared-memory parallelization of it using OpenMP. We also present an extensive experimental evaluation of our implementations using the LFR benchmark and real-world networks drawn from various domains. Compared with an implementation of LPA for community detection available in a widely used network analysis software, we achieve at most five times the F-Score while maintaining similar runtime for graphs with overlapping communities. We also compare DOLPA against an implementation of the Louvain method for community detection using the same LFR-graphs and show that DOLPA achieves about three times the F-Score at just 10% of the runtime. For connected component decomposition, our algorithm achieves orders of magnitude speedups over the basic LP-based algorithm on large-diameter graphs, up to 13.2× speedup over the Shiloach-Vishkin algorithm, and up to 1.6× speedup over Afforest on an Intel Xeon processor using 40 threads.

97 MATHEMATICS AND COMPUTING↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-Source Machine Learning and Thermoplastics Enhanced Aerostructure Manufacturing (mTEAM)

RTX Technology Research Center (RTRC), together with Collins Aerospace (Collins) and Oak Ridge National Laboratory (ORNL) has developed an Artificial Intelligence (AI) / Machine Learning (ML) guided solution to advance the manufacturing and assembly of high performance and lightweight thermoplastic composite (TPC) aerospace products. The solution aims to lower risk, cost and lead time for induction heating based welding and consolidation processes for TPC structure. The cost and lead time of part and material specific process development for induction welding (IW) and induction consolidation will be reduced by replacing traditional empirical methods with optimization methods that merge AI/ML and physics-based process simulations and process experiments with sensing and controls. TPC-IW process development is empirical in nature, and uncertainties in material & process behavior exist near & far from the induction coil. Physics-based simulations can be leveraged directly for process optimization but can be too computationally expensive to run in high fidelity and real time to do robust process optimization. The key impact of successful TPC induction consolidation and welding is cost & lead time reduction for part & material specific consolidation and welding recipes. This is an enabler for more rapid deployment of TPC structures via joining assembly, which can reduce energy & cost intensive usage of autoclaves & ovens. The solution aimed to advance the U.S. Department of Energy’s interests in using thermoplastics and automation in composite manufacturing for improvement of products for existing markets via increased production speeds, reduced costs, and lowered use of energy. Welded TPC structures can offer significant weight & energy savings for high-value commercial aerospace & industrial applications compared to metal & thermoset composite structures assembled by mechanical fastening and/or adhesive bonding. The project was organized into two Budget Periods. Budget Period 1 (BP1) was 15 months and its goal was to perform ML process optimization framework development & deployment on lab-coupon aerostructure components. A Go/No-Go Review was performed at the end of BP1 to verify fulfilment of key tasks & milestones to justify a Go Decision to move into the next Budget Period. Budget Period 2 (BP2) was 12 months and its goal was the deployment of the ML framework for ML process optimization of pilot industrial scale aerostructure components. The overall project aim was to develop & demonstrate ML-enhanced modeling framework that learns process-property mapping from multiple data sources at different fidelities. During BP1, the team accomplished key tasks & milestones to demonstrate the concept of multi-source ML for TPC aerostructure consolidation and assembly. First, the team completed documentation of induction based TPC heating requirements including baseline metrics to compare measured results against. Next the team completed demonstration of data generation from physics-based simulations for ML surrogate model generation and demonstrated the integration of physics-based simulation data into multi-source AI/ML algorithms. In parallel, the team established the lab-coupon scale induction welding system and completed a process to label and reduce generated data from physics-based simulation and experiments for ML surrogate models to enable multi-source ML model training & testing. To complete BP1, the team integrated physics-based simulation data and experimental data into multi-source ML algorithms. This was based on the team completing ML deployment of the induction welding on a lab system at RTRC and AI/ML deployment on existing induction welding line at Collins. ORNL visited both Collins and RTRC sites to witness the TPC induction welding process. Then, ORNL designed and constructed a new version of their vision-based sensing system better adapted to acquire process signals of the TPC induction welding process for process anomaly and defect detection. In BP2, the team accomplished key tasks & milestones to scale up multi-source ML for TPC aerostructure consolidation and assembly from the lab-coupon scale to the pilot-industrial scale. In BP2, the team demonstrated real time anomaly & defect detection via experiments performed by ORNL & RTRC. The team completed ML-optimization heating trials for TPC induction consolidation at Collins, and the team confirmed pilot industrial scale experimental data from Collins was compatible with the developed ML pipeline from RTRC. The team completed sub-element scale ML process optimization demonstration at RTRC, where the team leveraged RTRC’s robotic TPC welding setup to de-risk the ML process optimization by performing ML analysis of recorded temperatures to account for complex part features. Then, the team applied its ML-derived control strategies and ML process optimization framework at Collins to the pilot-industrial scale on a demo skin-stiffener part representative of a nacelle aerostructure fan cowl section. The key innovation is the AI/ML framework enabling effective process development of high performance, lightweight, energy efficient TPCs for composite aircraft structures.

36 MATERIALS SCIENCE↗

Artificial Intelligence for Multiphysics Nuclear Design Optimization with Additive Manufacturing

The geometric flexibility of additively manufactured metals and ceramics generates a very large and open design space that requires advanced modeling and simulation tools for physics simulations and the rigorous definition of design problems. This effort deploys artificial intelligence (AI) and machine learning (ML) algorithms to understand the design space, evaluate potential designs, and more efficiently generate optimized results. The Transformational Challenge Reactor (TCR) program is leveraging advances in several scientific areas—including materials, manufacturing, sensors and control systems, data analytics, and high-fidelity modeling and simulation—to accelerate the design, manufacturing, qualification, and deployment of advanced nuclear energy systems. Through a manufacturing-informed design approach, the TCR program seeks to integrate digital data for rapid nuclear innovation; accelerate the adoption of advances in manufacturing, materials, and computational sciences for nuclear applications; and dramatically reduce deployment costs and timelines for new nuclear reactor technologies. This report documents efforts under the TCR program to leverage advanced modeling and simulation techniques driven by AI/ML algorithms on high-performance computing (HPC) systems to yield more optimized TCR core designs. A multiphysics ML surrogate model was developed to run on the HPC architectures. The surrogate model is trained on high-fidelity simulation data of coupled neutronics and thermofluidics and is used to quickly evaluate thousands of candidate core designs in parallel, which drives the evolution of the cooling channel shapes to minimize temperature peaking and material stress. Outcomes from these activities provide design information and feedback into the core design efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Toward Large-Scale Image Segmentation on Summit

Semantic segmentation of images is an important computer vision task that emerges in a variety of application domains such as medical imaging, robotic vision and autonomous vehicles to name a few. While these domain-specific image analysis tasks involve relatively small image sizes (~ 10 2 × 10 2 ), there are many applications that need to train machine learning models on image data with extents that are orders of magnitude larger (~10 4 × 10 4 ). Training deep neural network (DNN) models on large extent images is extremely memory-intensive and often exceeds the memory limitations of a single graphical processing unit, a hardware accelerator of choice for computer vision workloads. Here, an efficient, sample parallel approach to train U-Net models on large extent image data sets is presented. Its advantages and limitations are analyzed and near-linear strong-scaling speedup demonstrated on 256 nodes (1536 GPUs) of the Summit supercomputer. Using a single node of the Summit supercomputer, an early evaluation of a recently released model parallel framework called GPipe is demonstrated to deliver ~ 2X speedup in executing a U-Net model with an order of magnitude larger number of trainable parameters than reported before. Performance bottlenecks for pipelined training of U-Net models are identified and mitigation strategies to improve the speedups are discussed. Together, these results open up the possibility of combining both approaches into a unified scalable pipelined and data parallel algorithm to efficiently train U-Net models with very large receptive fields on data sets of ultra-large extent images.

Seal, Sudip↗

Stabilization of the 81-channel coherent beam combination using machine learning

We develop a rapidly converging algorithm for stabilizing a large channel-count diffractive optical coherent beam combination. An 81-beam combiner is controlled by a novel, machine-learning based, iterative method to correct the optical phases, operating on an experimentally calibrated numerical model. A neural-network is trained to detect phase errors based on interference pattern recognition of uncombined beams adjacent to the combined one. Due to the non-uniqueness of solutions in the full space of possible phases, the network is trained within a limited phase perturbation/error range. This also reduces the number of samples needed for training. Simulations have proven that the network can converge in one step for small phase perturbations. When the trained neural-network is applied to a realistic case of 360 degree full range, an iterative scheme exploits random walking at the beginning, with the accuracy of prediction on phase feedback direction, to allow the neural-network to step into the training range for fast convergence. This neural-network-based iterative method of phase detection works tens of times faster than the commonly used stochastic parallel gradient descent approach (SPGD) using a single-detector and random dither when both are tested with random phase perturbations.

Wang, Dan↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Discovering causal structure with reproducing-kernel Hilbert space ε -machines

We merge computational mechanics’ definition of causal states (predictively equivalent histories) with reproducing-kernel Hilbert space (RKHS) representation inference. The result is a widely applicable method that infers causal structure directly from observations of a system’s behaviors whether they are over discrete or continuous events or time. A structural representation—a finite- or infinite-state kernel ϵ-machine—is extracted by a reduced-dimension transform that gives an efficient representation of causal states and their topology. In this way, the system dynamics are represented by a stochastic (ordinary or partial) differential equation that acts on causal states. We introduce an algorithm to estimate the associated evolution operator. Paralleling the Fokker–Planck equation, it efficiently evolves causal-state distributions and makes predictions in the original data space via an RKHS functional mapping. We demonstrate these techniques, together with their predictive abilities, on discrete-time, discrete-value infinite Markov-order processes generated by finite-state hidden Markov models with (i) finite or (ii) uncountably infinite causal states and (iii) continuous-time, continuous-value processes generated by thermally driven chaotic flows. The method robustly estimates causal structure in the presence of varying external and measurement noise levels and for very high-dimensional data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗