Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Designing alloys with process-mapping AI pre-trained on empirical knowledge

<span style="font-family: Calibri, sans-serif; font-size: 12pt;">Accelerated materials design should match the recent trends in the product development cycles. Materials data analytics can be used to significantly shorten development time of specialized alloys needed for next generation energy applications. However, it faces a challenge of scarce data available for training ML models. Incorporation of the domain knowledge into deep-learning graph structure via fuzzy pre-training and causal process imitation presents a viable approach to developing accurate data-driven models and reliable alloy design tools, with limited datasets. Artificial Intelligence (AI) was used in this study to incorporate such knowledge in the domain-specific computational tool, pyroMind. The tool provides not only novel design ideas but also their interpretation via physics and engineering concepts.</span>

Romanov, Vyacheslav↗

Domain knowledge-informed, process-mapping AI graph for designing Fe-based alloys

<span style="font-family: Calibri, sans-serif; font-size: 12pt;">Continuous improvement in efficiency of a power plant relies on designing materials for use at increasingly higher temperature and/or pressure, for 100,000s hours of operation. Due to complexity, non-linearity and high-dimensionality of the problem, traditional Machine Learning (ML) approaches require unreasonably large datasets for the data-driven model development. Science-based material and process engineering complements hard data with, sometimes soft and intuitive, empirical domain knowledge. Artificial Intelligence (AI) was used in this study to incorporate such knowledge into computational graph architecture (process-mimicking artificial neuron design, causal layer and graph structures, ensemble modeling of latent states) and learning procedures (variable transformation, fuzzy physics pre-training and freezing of deep layers, virtual microstructure representation, and adversarial multi-objective optimization). The first alloys design pathways suggested by the AI tool (pyroMind) passed a preliminary engineering review on soundness and transparency.</span>

Romanov, Vyacheslav↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Uncertainty quantification for Multiphase-CFD simulations of bubbly flows: a machine learning-based Bayesian approach supported by high-resolution experiments

In this paper, we developed a machine learning-based Bayesian approach to inversely quantify and reduce the uncertainties of multiphase computational fluid dynamics (MCFD) simulations for bubbly flows. The proposed approach is supported by high-resolution two-phase flow measurements, including those by double-sensor conductivity probes, high-speed imaging, and particle image velocimetry. Local distributions of key physical quantities of interest (QoIs), including the void fraction and phasic velocities, are obtained to support the Bayesian inference. In the process, the epistemic uncertainties of the closure relations are inversely quantified while the aleatory uncertainties from stochastic fluctuations of the system are evaluated based on experimental uncertainty analysis. The combined uncertainties are then propagated through the MCFD solver to obtain uncertainties of the QoIs, based on which probability-boxes are constructed for validation. The proposed approach relies on three machine learning methods: feedforward neural networks and principal component analysis for surrogate modeling, and Gaussian processes for model form uncertainty modeling. The whole process is implemented within the framework of an open-source deep learning library PyTorch with graphics processing unit (GPU) acceleration, thus ensuring the efficiency of the computation. The results demonstrate that with the support of high-resolution data, the uncertainties of MCFD simulations can be significantly reduced. The proposed approach has the potential for other applications that involve numerical models with empirical parameters.

42 ENGINEERING↗

Hierarchical deep reinforcement learning reveals a modular mechanism of cell movement

Time-lapse images of cells and tissues contain rich information about dynamic cell behaviours, which reflect the underlying processes of proliferation, differentiation and morphogenesis. However, we lack computational tools for effective inference. Here we exploit deep reinforcement learning (DRL) to infer cell–cell interactions and collective cell behaviours in tissue morphogenesis from three-dimensional (3D) time-lapse images. We use hierarchical DRL (HDRL), known for multiscale learning and data efficiency, to examine cell migrations based on images with a ubiquitous nuclear label and simple rules formulated from empirical statistics of the images. When applied to Caenorhabditis elegans embryogenesis, HDRL reveals a multiphase, modular organization of cell movement. Imaging with additional cellular markers confirms the modular organization as a novel migration mechanism, which we term sequential rosettes. Furthermore, HDRL forms a transferable model that successfully differentiates sequential rosettes-based migration from others. Our study demonstrates a powerful approach to infer the underlying biology from time-lapse imaging without prior knowledge.

59 BASIC BIOLOGICAL SCIENCES↗

Speeding up and reducing memory usage for scientific machine learning via mixed precision

Scientific machine learning (SciML) has emerged as a versatile approach to address complex computational science and engineering problems. Within this field, physics-informed neural networks (PINNs) and deep operator networks (DeepONets) stand out as the leading techniques for solving partial differential equations by incorporating both physical equations and experimental data. However, training PINNs and DeepONets require significant computational resources, including long computational times and large amounts of memory. In search of computational efficiency, training neural networks using half precision (float16) rather than the conventional single (float32) or double (float64) precision has gained substantial interest, given the inherent benefits of reduced computational time and memory consumed. However, we find that float16 cannot be applied to SciML methods, because of gradient divergence at the start of training, weight updates going to zero, and the inability to converge to a local minima. To overcome these limitations, we explore mixed precision, which is an approach that combines the float16 and float32 numerical formats to reduce memory usage and increase computational speed. Our experiments showcase that mixed precision training not only substantially decreases training times and memory demands but also maintains model accuracy. Here, we also reinforce our empirical observations with a theoretical analysis. The research has broad implications for SciML in various computational applications.

97 MATHEMATICS AND COMPUTING↗

Segmentation method comparison for residual fiber length measurement across tiled microscopy images

Fiber length distribution (FLD), in part, governs mechanical properties in discontinuous fiber composites, yet manual measurement methods limit the high-throughput characterization needed for materials design optimization. This study compares deep learning segmentation approaches for automated FLD measurement in large-field microscopy, evaluating how method choice affects the microstructural descriptors used in structure-property-processing relationships. A critical challenge is that high-resolution microscopy images (10,000×10,000 pixels) must be tiled for deep learning analysis, fragmenting fibers at boundaries. We demonstrate that segmentation method proves crucial for measurement accuracy. For example, instance segmentation with Slicing Aided Hyper Inference (SAHI) preserves individual fiber integrity across tiles while semantic segmentation prioritizes speed. Comparing against manual measurement of extracted carbon fibers, YOLOv11-SAHI matched manual ground truth (238 μm weighted mean) with 40x speedup (4.5 vs 167 minutes per image). U-Net provides rapid quantification although it is at the cost of reduced accuracy due only reliably measuring stand-alone fibers. Our comparative analysis reveals that instance segmentation with SAHI better preserves length measurements while semantic segmentation prioritizes speed, providing empirical guidance for method selection. The characterization provides essential inputs for mechanical property prediction models and inverse design workflows, accelerating composite materials development cycles.

Additive manufacturing↗

Multi-fidelity information fusion with concatenated neural networks

Recently, computational modeling has shifted towards the use of statistical inference, deep learning, and other data-driven modeling frameworks. Although this shift in modeling holds promise in many applications like design optimization and real-time control by lowering the computational burden, training deep learning models needs a huge amount of data. This big data is not always available for scientific problems and leads to poorly generalizable data-driven models. This gap can be furnished by leveraging information from physics-based models. Exploiting prior knowledge about the problem at hand, this study puts forth a physics-guided machine learning (PGML) approach to build more tailored, effective, and efficient surrogate models. For our analysis, without losing its generalizability and modularity, we focus on the development of predictive models for laminar and turbulent boundary layer flows. In particular, we combine the self-similarity solution and power-law velocity profile (low-fidelity models) with the noisy data obtained either from experiments or computational fluid dynamics simulations (high-fidelity models) through a concatenated neural network. We illustrate how the knowledge from these simplified models results in reducing uncertainties associated with deep learning models applied to boundary layer flow prediction problems. The proposed multi-fidelity information fusion framework produces physically consistent models that attempt to achieve better generalization than data-driven models obtained purely based on data. While we demonstrate our framework for a problem relevant to fluid mechanics, its workflow and principles can be adopted for many scientific problems where empirical, analytical, or simplified models are prevalent. In line with grand demands in novel PGML principles, this work builds a bridge between extensive physics-based theories and data-driven modeling paradigms and paves the way for using hybrid physics and machine learning modeling approaches for next-generation digital twin technologies.

42 ENGINEERING↗

Contrastive learning for robust representations of neutrino data

In neutrino physics, analyses often depend on large simulated datasets, making it essential for models to generalize effectively to real-world detector data. Contrastive learning, a well-established technique in deep learning, offers a promising solution to this challenge. By applying controlled data augmentations to simulated data, contrastive learning enables the extraction of robust and transferable features. This improves the ability of models trained on simulations to adapt to real experimental data distributions. In this paper, we investigate the application of contrastive learning methods in the context of neutrino physics. Through a combination of empirical evaluations and theoretical insights, we demonstrate how contrastive learning enhances model performance and adaptability. Additionally, we compare it to other domain adaptation techniques, highlighting the unique advantages of contrastive learning for this field. Published by the American Physical Society 2025

Wilkinson, Alex (ORCID:0000000253404506)↗

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

SympGNNs: Symplectic Graph Neural Networks for identifying high-dimensional Hamiltonian systems and node classification

Existing neural network models to learn Hamiltonian systems, such as SympNets, although accurate in low-dimensions, struggle to learn the correct dynamics for high-dimensional many-body systems. Herein, we introduce Symplectic Graph Neural Networks (SympGNNs) that can effectively handle system identification in high-dimensional Hamiltonian systems, as well as node classification. SympGNNs combine symplectic maps with permutation equivariance, a property of graph neural networks. Specifically, we propose two variants of SympGNNs: (i) G-SympGNN and (ii) LA-SympGNN, arising from different parameterizations of the kinetic and potential energy. We demonstrate the capabilities of SympGNN on two physical examples: a 40-particle coupled Harmonic oscillator, and a 2000-particle molecular dynamics simulation in a two-dimensional Lennard-Jones potential. Furthermore, we demonstrate the performance of SympGNN in the node classification task, achieving accuracy comparable to the state-of-the-art. Finally, we also empirically show that SympGNN can overcome the oversmoothing and heterophily problems, two key challenges in the field of graph neural networks.

Deep learning↗

Exploring Classification of Topological Priors With Machine Learning for Feature Extraction

In many scientific endeavors, increasingly abstract representations of data allow for new interpretive methodologies and conceptualization of phenomena. For example, moving from raw imaged pixels to segmented and reconstructed objects allows researchers new insights and means to direct their studies toward relevant areas. Thus, the development of new and improved methods for segmentation remains an active area of research. With advances in machine learning and neural networks, scientists have been focused on employing deep neural networks such as U-Net to obtain pixel-level segmentations, namely, defining associations between pixels and corresponding/referent objects and gathering those objects afterward. Topological analysis, such as the use of the Morse-Smale complex to encode regions of uniform gradient flow behavior, offers an alternative approach: first, create geometric priors, and then apply machine learning to classify. This approach is empirically motivated since phenomena of interest often appear as subsets of topological priors in many applications. Using topological elements not only reduces the learning space but also introduces the ability to use learnable geometries and connectivity to aid the classification of the segmentation target. Here, in this article, we describe an approach to creating learnable topological elements, explore the application of ML techniques to classification tasks in a number of areas, and demonstrate this approach as a viable alternative to pixel-level classification, with similar accuracy, improved execution time, and requiring marginal training data.

97 MATHEMATICS AND COMPUTING↗

Deep potential molecular dynamics simulations of ion-enhanced etching of silicon by atomic chlorine

The continued development of plasma-assisted processing techniques requires a fundamental understanding of plasma-surface interactions. Molecular dynamics (MD) simulations have been employed to complement experimental studies and better understand the properties of such systems. Recently, machine learning (ML) methods have enabled the development of ab initio-based interatomic potentials, which can be generalized to complex combinations of multiple atom types. In this work, we use ML potentials developed using the Deep Potential Molecular Dynamics (DeepMD) framework to provide a model of ion-enhanced etching of Si by Cl atoms. We demonstrate the importance of proper selection of the training data set to the accuracy of the DeepMD model and compare our results to MD results using empirical potentials, as well as to experimental measurements. Exposure of undoped Si at 300 K to thermal Cl atoms yields a steady-state Cl coverage of 1.25 monolayers, which is slightly lower than the value obtained in previous experimental studies. Predictions of Si etch yields by simultaneous Cl atom and Ar + ion impacts as a function of ion energy, neutral to ion flux ratio, and angle of incidence of the ions are in reasonably good agreement with classical MD results and experimental measurements. Finally, etch yields and SiCl x mixed layer thicknesses during simultaneous bombardment of the Si(100) surface by Cl atoms and Cl + ions are in good agreement with experimental data. In conclusion, the present work is a necessary condition for the extension of the DeepMD procedure to more complex systems of interest in plasma-surface interactions.

Artificial neural networks↗

Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics

Conventional machine-learning (ML) models in computational chemistry learn to directly predict molecular properties using quantum chemistry only for reference data. While these heuristic ML methods show quantum-level accuracy with speeds several orders of magnitude faster than traditional quantum chemistry methods, they suffer from poor extensibility and transferability; i.e., their accuracy degrades on large or new chemical systems. Incorporating quantum chemistry frameworks into the ML models directly solves this problem. Here we take the structure of semiempirical quantum mechanics (SEQM) methods to construct dynamically responsive Hamiltonians. SEQM methods use empirical parameters fitted to experimental properties to construct reduced-order Hamiltonians, facilitating much faster calculations than ab initio methods but with compromised accuracy. By replacing these static parameters with machine-learned dynamic values inferred from the local environment, we greatly improve the accuracy of the SEQM methods. Trained on molecular energies and atomic forces, these dynamically generated Hamiltonian parameters show a strong correlation with atomic hybridization and bonding. Trained with only about 60,000 small organic molecular conformers, the resulting model retains interpretability, extensibility, and transferability when testing on much larger chemical systems and predicting various molecular properties. Overall, this work demonstrates the virtues of incorporating physics-based descriptions with ML to develop models that are simultaneously accurate, transferable, and interpretable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Novel LDPP-MADDPG Approach for Distributed Power Allocation in mmWave Cellular Networks

This paper considers the problem of distributed beam scheduling and power allocation problem in millimeter- Wave (mmWave) cellular networks, in which multiple Base Stations (BSs) operate as individual operators over a shared spectrum. We propose a novel learning-aided approach that integrates the Lyapunov Drift-Plus-Penalty (LDPP) framework and Multi-agent Deep Deterministic Policy Gradient (MADDPG) reinforcement learning algorithms. This offers a powerful approach to learning stable and constraint-aware policies, reaping the joint benefit of both LDPP and MADDPG, in complex multiagent environments. The major challenge for this approach is to integrate these two approaches in a meaningful and effective manner. The key idea to solve this problem is to introduce a novel feature of local observation that incorporates potential negative value of the reward function due to the stochastic constraints introduced by the LDPP framework. Empirical results demonstrate that our proposed scheme outperforms the baseline methods under various conditions.

99 - GENERAL AND MISCELLANEOUS↗

DeepGraphONet: A Deep Graph Operator Network to Learn and Zero-Shot Transfer the Dynamic Response of Networked Systems

This article develops a deep graph operator network (DeepGraphONet) framework that learns to approximate the dynamics of a complex system (e.g., the power grid or traffic) with an underlying subgraph structure. Here, we build our DeepGraphONet by fusing the ability of graph neural networks to exploit spatially correlated graph information and deep operator networks to approximate the solution operator of dynamical systems. The resulting DeepGraphONet can then predict the dynamics within a given short/medium-term time horizon by observing a finite history of the graph state information. Furthermore, we design our DeepGraphONet to be resolution independent. That is, we do not require the finite history to be collected at the exact/same resolution. In addition, to disseminate the results from a trained DeepGraphONet, we design a zero-shot learning strategy that enables using it on a different subgraph. Finally, empirical results on the transient stability prediction problem of power grids and traffic flow forecasting problem of a vehicular system illustrate the effectiveness of the proposed DeepGraphONet.

24 POWER TRANSMISSION AND DISTRIBUTION↗

“Understanding Robustness Lottery”: A Geometric Visual Comparative Analysis of Neural Network Pruning Approaches

Deep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. Here, this work aims to shed light on how different pruning methods alter the network’s internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption.

Li, Zhimin [Univ. of Utah, Salt Lake City, UT (Uni↗

Z-Sequence: photometric redshift predictions for galaxy clusters with sequential random k-nearest neighbours

ABSTRACT We introduce Z-Sequence, a novel empirical model that utilizes photometric measurements of observed galaxies within a specified search radius to estimate the photometric redshift of galaxy clusters. Z-Sequence itself is composed of a machine learning ensemble based on the k-nearest neighbours algorithm. We implement an automated feature selection strategy that iteratively determines appropriate combinations of filters and colours to minimize photometric redshift prediction error. We intend for Z-Sequence to be a standalone technique but it can be combined with cluster finders that do not intrinsically predict redshift, such as our own DEEP-CEE. In this proof-of-concept study, we train, fine-tune, and test Z-Sequence on publicly available cluster catalogues derived from the Sloan Digital Sky Survey. We determine the photometric redshift prediction error of Z-Sequence via the median value of |Δ$z$|/(1 + $z$) (across a photometric redshift range of 0.05 ≤ $z$ ≤ 0.6) to be ∼0.01 when applying a small search radius. The photometric redshift prediction error for test samples increases by 30–50 per cent when the search radius is enlarged, likely due to line-of-sight interloping galaxies. Eventually, we aim to apply Z-Sequence to upcoming imaging surveys such as the Legacy Survey of Space and Time to provide photometric redshift estimates for large samples of as yet undiscovered and distant clusters.

Chan, Matthew C.↗