Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Feature Space”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

SA-GAT-SR: self-adaptable graph attention networks with symbolic regression for high-fidelity material property prediction

Recent advances in machine learning have demonstrated an enormous utility of deep learning approaches, particularly Graph Neural Networks (GNNs) for materials science. These methods have emerged as powerful tools for high-throughput prediction of material properties, offering a compelling enhancement and alternative to traditional first-principles calculations. While the community has predominantly focused on developing increasingly complex and universal models to enhance predictive accuracy, such approaches often lack physical interpretability and insights into materials behavior. Here, we introduce a novel computational paradigm—Self-Adaptable Graph Attention Networks integrated with Symbolic Regression (SA-GAT-SR)—that synergistically combines the predictive capability of GNNs with the interpretative power of symbolic regression. Our framework employs a self-adaptable encoding algorithm that automatically identifies and adjust attention weights so as to screen critical features from an expansive 180-dimensional feature space while maintaining O(n) computational scaling. The integrated SR module subsequently distills these features into compact analytical expressions that explicitly reveal quantum-mechanically meaningful relationships, achieving 23 × acceleration compared to conventional SR implementations that heavily rely on first-principle calculations-derived features as input. This work suggests a new framework in computational materials science, bridging the gap between predictive accuracy and physical interpretability, offering valuable physical insights into material behavior.

36 MATERIALS SCIENCE↗

ASDFL: An adaptive super‐pixel discriminative feature‐selective learning for vehicle matching

Abstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method.

Qin, Rong↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Reduced‐Order Modeling of Energetic Materials Using Physics‐Aware Recurrent Convolutional Neural Networks in a Latent Space (LatentPARC)

Physics-aware deep learning (PADL) has gained popularity for use in spatiotemporal dynamics simulations, such as those in computational modeling of energetic materials (EM). We show that the challenge PADL methods face while learning complex field evolution problems can be simplified and accelerated by decoupling it into two tasks: learning complex geometric features in evolving fields and modeling dynamics over these features in a lower-dimensional feature space. We build upon our previous work on physics-aware recurrent convolutional neural networks (PARC). PARC embeds knowledge of underlying physics into its neural network architecture for more robust and accurate prediction of evolving physical fields. PARC was shown to effectively learn complex nonlinear features such as the formation of hotspots and coupled shock fronts in various initiation scenarios of EMs, as a function of microstructures, serving effectively as a microstructure-aware burn model. Here, we further accelerate PARC and reduce its computational cost by projecting the original dynamics onto a lower-dimensional invariant manifold, or “latent space.” The projected latent representation encodes the complex geometry of evolving fields (e.g., temperature and pressure) in a set of data-driven features. The reduced dimension of this latent space allows us to learn the dynamics during the initiation of EM with a lighter and more efficient model. We observe a significant decrease in training and inference time while maintaining results comparable to PARC at inference. This work takes steps towards enabling rapid prediction of EM thermomechanics at larger scales and characterization of EM structure–property–performance linkages at a full application scale.

Mathematics and Computing↗

Inverse design of pore wall chemistry and topology through active learning of surface group interactions

Design of next-generation membranes requires a nanoscopic understanding of the effect of biologically inspired heterogeneous surface chemistries and topologies (roughness) on local water and solute behavior. In particular, the rejection of small, neutral solutes, such as boric acid, poses a heretofore unsolved challenge. In prior work, a computational inverse design technique using an evolutionary optimization successfully uncovered new surface design strategies for optimized transport of water over solutes in smooth, model pores consisting of two surface chemistries. However, extending such an approach to more complex (and realistic) scenarios involving many surface chemistries as well as surface roughness is challenging due to the expanded design space. In this work, we develop a new approach that uses active learning to optimize in a reduced feature space of surface group interactions, finding parameters that lead to their assembly into ordered, optimal patterns. This approach rapidly identifies novel surface functionalizations that maximize the difference in water and boric acid transport through the nanopore. Moreover, we find that the roughness of the nanopore wall, independent of its chemistry, can be leveraged to enhance transport selectivity: oscillations in the pore wall diameter optimally inhibit boric acid transport by creating energetic wells from which the solute must escape to transport down the pore. Furthermore, this proof-of-concept demonstrates the potential for active learning strategies, in concert with molecular simulations, to rapidly navigate complex design spaces of aqueous interfaces and is promising as a tool for engineering water-mediated surface interactions for a broad range of applications.

36 MATERIALS SCIENCE↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗

Predicting industrial building energy consumption with statistical and machine-learning models informed by physical system parameters

The industrial sector consumes about one-third of global energy, making them a frequent target for energy use reduction. Variation in energy usage is observed with weather conditions, as space conditioning needs to change seasonally, and with production, energy-using equipment is directly tied to production rate. Previous models were based on engineering analyses of equipment and relied on site-specific details. Others consisted of single-variable regressors that did not capture all contributions to energy consumption. Further, new modeling techniques could be applied to rectify these weaknesses. Applying data from 45 different manufacturing plants obtained from industrial energy audits, a supervised machine-learning model is developed to create a general predictor for industrial building energy consumption. The model uses features of air enthalpy, solar radiation, and wind speed to predict weather-dependency; motor, steam, and compressed air system parameters to capture support equipment contributions; and operating schedule, production rate, number of employees, and floor area to determine production-dependency. Results showed that a model that used a linear regressor over a transformed feature space could outperform a support vector machine and utilize features more representative of physical systems. Using informed parameters to build a reliable predictor will more accurately characterize a manufacturing facility's energy savings opportunities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Learning with Delayed Rewards—A Case Study on Inverse Defect Design in 2D Materials

Defect dynamics in materials are of central importance to a broad range of technologies from catalysis to energy storage systems to microelectronics. Material functionality depends strongly on the nature and organization of defects-their arrangements often involve intermediate or transient states that present a high barrier for transformation. The lack of knowledge of these intermediate states and the presence of this energy barrier presents a serious challenge for inverse defect design, especially for gradient-based approaches. Here, we present a reinforcement learning (RL) [Monte Carlo Tree Search (MCTS)] based on delayed rewards that allow for efficient search of the defect configurational space and allows us to identify optimal defect arrangements in low-dimensional materials. Using a representative case of two-dimensional MoS 2 , we demonstrate that the use of delayed rewards allows us to efficiently sample the defect configurational space and overcome the energy barrier for a wide range of defect concentrations (from 1.5 to 8% S vacancies)-the system evolves from an initial randomly distributed S vacancies to one with extended S line defects consistent with previous experimental studies. Detailed analysis in the feature space allows us to identify the optimal pathways for this defect transformation and arrangement. Comparison with other global optimization schemes like genetic algorithms suggests that the MCTS with delayed rewards takes fewer evaluations and arrives at a better quality of the solution. The implications of the various sampled defect configurations on the 2H to 1T phase transitions in MoS 2 are discussed. In this study, we introduce a RL strategy employing delayed rewards that can accelerate the inverse design of defects in materials for achieving targeted functionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tuning 3-D Nanomaterial Architectures Using Atomic Layer Deposition to Direct Solution Synthesis

The ability to synthesize nanoarchitected materials with tunable geometries provides a means to control their functional properties, with applications in biological, environmental, and energy fields. To this end, various bottom-up and top-down synthesis processes have been developed. However, many of these processes require prepatterning or etching steps, making them challenging to scale-up to complex, nonplanar substrates. Furthermore, the ability to integrate nanomaterials into hierarchical arrays with precise control of feature spacing and orientation remains a challenge. One approach to overcome these patterning challenges is the use of surface modification layers to guide the resulting geometry of nanomaterial architectures grown from the substrate. A powerful strategy to accomplish this is what we will refer to as “surface-directed assembly,” where the resulting geometric parameters (feature size, shape, orientation) are predetermined by the initial surface layer. In particular, the use of Atomic Layer Deposition (ALD) to form a surface layer, followed by solution-based growth processes, has the ability to synthesize architected structures with tunable geometries on complex, nonplanar surfaces. Over the past decade, we have reported a series of studies where surface-directed assembly is used to synthesize ZnO nanowires (NWs) on top of a variety of substrates. In this case, a thin film of ZnO is deposited onto the substrate using ALD, which can guide the NW diameter, spacing, and angular orientation with respect to the substrate by controlling epitaxial relationships. Furthermore, we have shown that by depositing a submonolayer overcoat of a secondary material (e.g., amorphous TiO 2 ), nucleation sites are partially blocked, which can further tune the spacing between nanowires while minimizing changes to their other geometric properties. This approach can be used to generate multilevel hierarchical structures, such as hyperbranched NW arrays with tunable control of each level of hierarchy using ALD. Finally, we have demonstrated that the tunable control of geometric parameters can be scaled-up to curved, nonplanar substrates. This highlights the power of ALD to conformally and uniformly deposit the seed layers on complex substrates with subnanometer precision. To complement these seeded hydrothermal approaches, we expanded this strategy to include conversion chemistry of the initial ALD seed layers. For example, by replacing ZnO with Al 2 O 3 as the seed layer without changing the hydrothermal growth conditions, Al–Zn layered-double hydroxide nanosheets can be formed instead of nanowires. In another example of conversion chemistry, a solution anion-exchange process was used to incorporate sulfur into ALD metal oxide films. In both of these conversion processes, the properties of the initial ALD film enabled tuning of the resulting nanostructure geometry. In this Account, we describe the use of ALD to guide the growth of diverse nanomaterial systems, with tunable control over their geometry and composition. We further show how these approaches can be used to tune functional properties for a range of applications, including superomniphobic surfaces, antibiofouling coatings, and photocatalysis. In conclusion, we conclude with an outlook on how the combination of ALD and solution synthesis can enable future directions in scalable nanomanufacturing to overcome the limitations of traditional top-down and bottom-up approaches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Non-parametric Lagrangian biasing from the insights of neural nets

We present a Lagrangian model of galaxy clustering bias in which we train a neural net using the local properties of the smoothed initial density field to predict the late-time mass-weighted halo field.By fitting the mass-weighted halo field in the AbacusSummit simulations at z = 0.5, we find that including three coarsely spaced smoothing scales gives the best recovery of the halo power spectrum. Adding more smoothing scales may lead to 2–5% underestimation of the large-scale power and can cause the neural net to overfit.We find that the fitted halo-to-mass ratio can be well described by two directions in the original high-dimension feature space.Projecting the original features into these two principal components and re-training the neural net either reproduces the original training result, or outperforms it with a better match of the halo power spectrum. The elements of the principal components are unlikely to be assigned physical meanings, partly owing to the features being highly correlated between different smoothing scales.Our work illustrates a potential need to include multiple smoothing scales when studying galaxy bias, and this can be done easily with machine-learning methods that can take in high dimensional input feature space.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating particle-in-cell kinetic plasma simulations via reduced-order modeling of space-charge dynamics using dynamic mode decomposition

We present a data-driven reduced-order modeling of the space-charge dynamics for electromagnetic particle-in-cell (EMPIC) plasma simulations based on dynamic mode decomposition (DMD). The dynamics of the charged particles in kinetic plasma simulations such as EMPIC is manifested through the plasma current density defined along the edges of the spatial mesh. We showcase the efficacy of DMD in modeling the time evolution of current density through a low-dimensional feature space. Not only do such DMD based predictive reduced-order models help accelerate EMPIC simulations, they also have the potential to facilitate investigative analysis and control applications. Here, we demonstrate the proposed DMD-EMPIC scheme for reduced-order modeling of current density and speedup in EMPIC simulations involving electron beam under the influence of magnetic field, virtual cathode oscillations, and backward wave oscillator.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Merger identification through photometric bands, colours, and their errors

Aims. We present the application of a fully connected neural network (NN) for galaxy merger identification using exclusively photometric information. Our purpose is not only to test the method’s efficiency, but also to understand what merger properties the NN can learn and what their physical interpretation is. Methods. We created a class-balanced training dataset of 5860 galaxies split into mergers and non-mergers. The galaxy observations came from SDSS DR6 and were visually identified in Galaxy Zoo. The 2930 mergers were selected from known SDSS mergers and the respective non-mergers were the closest match in both redshift and r magnitude. The NN architecture was built by testing a different number of layers with different sizes and variations of the dropout rate. We compared input spaces constructed using: the five SDSS filters: u, g, r, i, and z; combinations of bands, colours, and their errors; six magnitude types; and variations of input normalization. Results. We find that the fibre magnitude errors contribute the most to the training accuracy. Studying the parameters from which they are calculated, we show that the input space built from the sky error background in the five SDSS bands alone leads to 92.64 ± 0.15% training accuracy. We also find that the input normalization, that is to say, how the data are presented to the NN, has a significant effect on the training performance. Conclusions. We conclude that, from all the SDSS photometric information, the sky error background is the most sensitive to merging processes. This finding is supported by an analysis of its five-band feature space by means of data visualization. Moreover, studying the plane of the g and r sky error bands shows that a decision boundary line is enough to achieve an accuracy of 91.59%.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven model for divertor plasma detachment prediction

We present a fast and accurate data-driven surrogate model for divertor plasma detachment prediction leveraging the latent feature space concept in machine learning research. Our approach involves constructing and training two neural networks: an autoencoder that finds a proper latent space representation (LSR) of plasma state by compressing the multi-modal diagnostic measurements and a forward model using multi-layer perception (MLP) that projects a set of plasma control parameters to its corresponding LSR. By combining the forward model and the decoder network from autoencoder, this new data-driven surrogate model is able to predict a consistent set of diagnostic measurements based on a few plasma control parameters. In order to ensure that the crucial detachment physics is correctly captured, highly efficient 1D UEDGE model is used to generate training and validation data in this study. The benchmark between the data-driven surrogate model and UEDGE simulations shows that our surrogate model is capable of providing accurate detachment prediction (usually within a few per cent relative error margin) but with at least four orders of magnitude speed-up, indicating that performance-wise, it has the potential to facilitate integrated tokamak design and plasma control. Comparing with the widely used two-point model and/or two-point model formatting, the new data-driven model features additional detachment front prediction and can be easily extended to incorporate richer physics. This study demonstrates that the complicated divertor and scrape-off-layer plasma state has a low-dimensional representation in latent space. Understanding plasma dynamics in latent space and utilising this knowledge could open a new path for plasma control in magnetic fusion energy research.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Explicit physics-informed neural networks for nonlinear closure: The case of transport in tissues

In upscaling methods, closures for nonlinear problems present a well-known challenge. While a number of theoretical methods have been proposed for handling such closures, nonlinearities still remain a significant obstacle for many problems. In this work, we use a combination of formal upscaling and data-driven machine learning for explicitly closing a nonlinear transport and reaction process in multiscale tissues. The classical effectiveness factor model is used to formulate the macroscale reaction kinetics. We train a multilayer perceptron network using training data generated by direct numerical simulations over microscale examples. Once trained, the network is used in an algorithm for numerically solving the upscaled (coarse-grained) differential equation describing mass transport and reaction in two example tissues. The network is described as being explicit in the sense that the network is trained using macroscale concentrations and gradients of concentration as components of the feature space rather than incorporating them as part of a constraint in the optimization process. Network training and solutions to the macroscale transport equations were computed for two different tissues. The two tissue types (brain and liver) exhibit markedly different geometrical complexity and spatial scale (cell size and sample size). The upscaled solutions for the average concentration are compared with numerical solutions derived from the microscale concentration fields by a posteriori averaging. There are three outcomes of this work of particular note. 1) Our overall approach results in an upscaled nonlinear PDE. The PDE is closed using a neural network, and our approach results in the definition of the classical effectiveness factor for effecting closure. 2) We identify particular source terms for the closure problem that are important for representing the structure of the closure. These source terms involve macroscale concentrations and their gradients. We adopt these source terms to use as explicit features in the learning algorithm. We find the trained networks that include the macroscale source terms generate models that are able to predict the correction factor with increased fidelity over those that do not. 3) We find that the trained network exhibits good generalizability, and it is able to predict the effectiveness factor with high fidelity for realistically-structured tissues despite the significantly different scale and geometrical complexity of the two example tissue types. This latter result emphasizes our purposeful connection between conventional averaging methods with the use of machine learning for closure; this contrasts with some machine learning methods for upscaling where the exact form of the macroscale equation remains unknown.

97 MATHEMATICS AND COMPUTING↗

Nonlocal Kernel Network (NKN): a Stable and Resolution-Independent Deep Neural Network.

Neural operators have recently become popular tools for designing solution maps between function spaces in the form of neural networks. Differently from classical scientific machine learning approaches that learn parameters of a known partial differential equation (PDE) for a single instance of the input parameters at a fixed resolution, neural operators approximate the solution map of a family of PDEs [6, 7]. Despite their success, the uses of neural operators are so far restricted to relatively shallow neural networks and confined to learning hidden governing laws. In this work, we propose a novel nonlocal neural operator, which we refer to as nonlocal kernel network (NKN), that is resolution independent, characterized by deep neural networks, and capable of handling a variety of tasks such as learning governing equations and classifying images. Our NKN stems from the interpretation of the neural network as a discrete nonlocal diffusion reaction equation that, in the limit of infinite layers, is equivalent to a parabolic nonlocal equation, whose stability is analyzed via nonlocal vector calculus. The resemblance with integral forms of neural operators allows NKNs to capture long-range dependencies in the feature space, while the continuous treatment of node-to-node interactions makes NKNs resolution independent. The resemblance with neural ODEs, reinterpreted in a nonlocal sense, and the stable network dynamics between layers allow for generalization of NKN’s optimal parameters from shallow to deep networks. This fact enables the use of shallow-to-deep initialization techniques [8]. Our tests show that NKNs outperform baseline methods in both learning governing equations and image classification tasks and generalize well to different resolutions and depths.

97 MATHEMATICS AND COMPUTING↗