Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multimodal deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Using Neural Architecture Search for Improving Software Flaw Detection in Multimodal Deep Learning Models

Software flaw detection using multimodal deep learning models has been demonstrated as a very competitive approach on benchmark problems. In this work, we demonstrate that even better performance can be achieved using neural architecture search (NAS) combined with multimodal learning models. We adapt a NAS framework aimed at investigating image classification to the problem of software flaw detection and demonstrate improved results on the Juliet Test Suite, a popular benchmarking data set for measuring performance of machine learning models in this problem domain.

97 MATHEMATICS AND COMPUTING↗

Using Neural Architecture Search for Improving Software Flaw Detection in Multimodal Deep Learning Models

Software flaw detection using multimodal deep learning models has been demonstrated as a very competitive approach on benchmark problems. In this work, we demonstrate that even better performance can be achieved using neural architecture search (NAS) combined with multimodal learning models. We adapt a NAS framework aimed at investigating image classification to the problem of software flaw detection and demonstrate improved results on the Juliet Test Suite, a popular benchmarking data set for measuring performance of machine learning models in this problem domain.

97 MATHEMATICS AND COMPUTING↗

Multimodal Deep Learning for Flaw Detection in Software Programs

We explore the use of multiple deep learning models for detecting flaws in software programs. Current, standard approaches for flaw detection rely on a single representation of a software program (e.g., source code or a program binary). We illustrate that, by using techniques from multimodal deep learning, we can simultaneously leverage multiple representations of software programs to improve flaw detection over single representation analyses. Specifically, we adapt three deep learning models from the multimodal learning literature for use in flaw detection and demonstrate how these models outperform traditional deep learning models. We present results on detecting software flaws using the Juliet Test Suite and Linux Kernel.

97 MATHEMATICS AND COMPUTING↗

Automated Grain Boundary (GB) Segmentation and Microstructural Analysis in 347H Stainless Steel Using Deep Learning and Multimodal Microscopy

Austenitic 347H stainless steel offers superior mechanical properties and corrosion resistance required for extreme operating conditions such as high temperature. The change in microstructure due to composition and process variations is expected to impact material properties. Identifying microstructural features such as grain boundaries thus becomes an important task in the process-microstructure-properties loop. Applying convolutional neural network (CNN)-based deep learning models is a powerful technique to detect features from material micrographs in an automated manner. In contrast to microstructural classification, supervised CNN models for segmentation tasks require pixel-wise annotation labels. However, manual labeling of the images for the segmentation task poses a major bottleneck for generating training data and labels in a reliable and reproducible way within a reasonable timeframe. Microstructural characterization especially needs to be expedited for faster material discovery by changing alloy compositions. Here, in this study, we attempt to overcome such limitations by utilizing multimodal microscopy to generate labels directly instead of manual labeling. We combine scanning electron microscopy images of 347H stainless steel as training data and electron backscatter diffraction micrographs as pixel-wise labels for grain boundary detection as a semantic segmentation task. The viability of our method is evaluated by considering a set of deep CNN architectures. We demonstrate that despite producing instrumentation drift during data collection between two modes of microscopy, this method performs comparably to similar segmentation tasks that used manual labeling. Additionally, we find that naïve pixel-wise segmentation results in small gaps and missing boundaries in the predicted grain boundary map. By incorporating topological information during model training, the connectivity of the grain boundary network and segmentation performance is improved. Finally, our approach is validated by accurate computation on downstream tasks of predicting the underlying grain morphology distributions which are the ultimate quantities of interest for microstructural characterization.

36 MATERIALS SCIENCE↗

Deep Learning on Multimodal Chemical and Whole Slide Imaging Data for Predicting Prostate Cancer Directly from Tissue Images

Prostate cancer is one of the most common cancers globally and is the second most common cancer in the male population in the US. Here we develop a study based on correlating the hematoxylin and eosin (H&E)-stained biopsy data with MALDI mass-spectrometric imaging data of the corresponding tissue to determine the cancerous regions and their unique chemical signatures and variations of the predicted regions with original pathological annotations. We obtain features from high-resolution optical micrographs of whole slide H&E stained data through deep learning and spatially register them with mass spectrometry imaging (MSI) data to correlate the chemical signature with the tissue anatomy of the data. We then use the learned correlation to predict prostate cancer from observed H&E images using trained coregistered MSI data. This multimodal approach can predict cancerous regions with ~80% accuracy, which indicates a correlation between optical H&E features and chemical information found in MSI. Further, we show that such paired multimodal data can be used for training feature extraction networks on H&E data which bypasses the need to acquire expensive MSI data and eliminates the need for manual annotation saving valuable time. Two chemical biomarkers were also found to be predicting the ground truth cancerous regions. This study shows promise in generating improved patient treatment trajectories by predicting prostate cancer directly from readily available H&E-stained biopsy images aided by coregistered MSI data.

60 APPLIED LIFE SCIENCES↗

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin↗

Joint Analysis of Program Data Representations using Machine Learning for Improved Software Assurance and Development Capabilities

We explore the use of multiple deep learning models for detecting flaws in software programs. Current, standard approaches for flaw detection rely on a single representation of a software program (e.g., source code or a program binary). We illustrate that, by using techniques from multimodal deep learning, we can simultaneously leverage multiple representations of software programs to improve flaw detection over single representation analyses. Specifically, we adapt three deep learning models from the multimodal learning literature for use in flaw detection and demonstrate how these models outperform traditional deep learning models. We present results on detecting software flaws using the Juliet Test Suite and Linux Kernel.

97 MATHEMATICS AND COMPUTING↗

AI-Based Analytics and Energy Modeling Framework for Characterizing Urban Energy Systems

Developing location-specific district energy models is essential for understanding energy patterns and supporting efficient management and planning decisions. However, accurately characterizing these models remains challenging due to gaps in building characteristics and labor-intensive traditional modeling workflows. To address these challenges, we develop an AI-based framework that integrates top-down and bottom-up building energy data to automate urban energy model characterization. The framework trains multimodal deep learning models using heterogeneous ResStockTM datasets to infer missing building characteristics from varying levels of known information and generate simulation-ready inputs for district-scale energy modeling. It also employs a conditioning-based injection approach to generate ”what-if” scenarios, enabling users to explore retrofit, efficiency, and technology-upgrade pathways. Integrated within URBANoptTM, a bottom-up district energy modeling platform for simulating co-located buildings, the framework infers detailed building-level inputs required for bottom-up simulations. Both localized and generalized AI models are developed to learn relationships across categorical, numerical, and time-series data, enabling reconstruction of missing attributes and generation of targeted upgrade scenarios. We demonstrate this methodology on a residential neighborhood in Baltimore, MD, assessing internal consistency against ResStock reference data and URBANopt simulation, and comparing selected attributes against real-world building characteristics. Results show strong overall predictive accuracy in data completion and scenario generation, with localized and generalized models offering complementary trade-offs between precision and scalability. Overall, our automated framework streamlines energy modeling and provides a reliable framework for urban building energy characterization.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Investigating permafrost carbon dynamics in Alaska with artificial intelligence

Abstract Positive feedbacks between permafrost degradation and the release of soil carbon into the atmosphere impact land–atmosphere interactions, disrupt the global carbon cycle, and accelerate climate change. The widespread distribution of thawing permafrost is causing a cascade of geophysical and biochemical disturbances with global impacts. Currently, few earth system models account for permafrost carbon feedback (PCF) mechanisms. This research study integrates artificial intelligence (AI) tools and information derived from field-scale surveys across the tundra and boreal landscapes in Alaska. We identify and interpret the permafrost carbon cycling links and feedback sensitivities with GeoCryoAI, a hybridized multimodal deep learning (DL) architecture of stacked convolutionally layered, memory-encoded recurrent neural networks (NN). This framework integratesin-situmeasurements and flux tower observations for teacher forcing and model training. Preliminary experiments to quantify, validate, and forecast permafrost degradation and carbon efflux across Alaska demonstrate the fidelity of this data-driven architecture. More specifically, GeoCryoAI logs the ecological memory and effectively learns covariate dynamics while demonstrating an aptitude to simulate and forecast PCF dynamics—active layer thickness (ALT), carbon dioxide flux (CO 2 ), and methane flux (CH 4 )—with high precision and minimal loss (i.e. ALT RMSE : 1.327 cm [1969–2022]; CO 2 RMSE : 0.697µmolCO 2 m −2 s −1 [2003–2021]; CH 4 RMSE : 0.715 nmolCH 4 m −2 s −1 [2011–2022]). ALT variability is a sensitive harbinger of change, a unique signal characterizing the PCF, and our model is the first characterization of these dynamics across space and time.

Environmental Sciences & Ecology↗

Multimodal Data Representation with Deep Learning for Extracting Cancer Characteristics from Clinical Text

This paper presents a multimodal data representation to improve the performance of deep learning models for extracting cancer key characteristics from unstructured text in pathology reports. Specifically, in addition to using the text as the input to deep learning models, we use concept unique identifiers (CUIs) as another source of information to the models. We analyze the performance of different text and CUI data representations, including word embeddings and bag of embeddings (BOE), with a convolutional neural network (CNN) and a fully connected multilayer perceptron neural network (MLP-NN). The high level document embeddings from text and CUI inputs are combined by concatenating them and then applying a classifier. The model is used for extracting cancer subsite and histology from pathology reports. These two classification tasks have a large number of labels, i.e. 317 for subsite and 556 for histology, with extreme class imbalance. We compare the performance of the developed DL models across the two tasks based on micro- and macro-F1 scores. The evaluation shows that a multi-channel DL model that utilizes text represented by word embeddings and CUIs represented by BOE outperforms other DL models. Also, this approach significantly improves the model performance on low prevalence classes.

Alawad, Mohammed↗

Bridging multimodal microscopy for advanced characterization on nuclear fuel using machine learning

Uranium dioxide (UO 2 ), widely used as driver fuel in light water reactors, experiences microstructure and property change by nuclear fission reactions. This paper bridges the characterization of fresh UO 2 fuel at different length scales, serving as a baseline for future post irradiation examination of irradiated UO 2 fuel. To characterize the microstructural change of nuclear fuel, modern approaches cover a wide range of length scales through different characterization techniques, such as mm scale for Synchrotron-based X-ray computed tomography (SXCT) and microscale for focused ion beam (FIB) and scanning electron microscopy (SEM). It is challenging to bridge the data and knowledge of the same sample in different length scales. This paper proposed a deep learning framework leveraging transfer learning to detect microstructural defects, trained from a sparse FIB, SEM, and SXCT images. The proposed model achieved superior performance in defect segmentation on multiscale microscopic data compared to four of the latest deep learning models.

36 MATERIALS SCIENCE↗

Data augmentation and multimodal learning for predicting drug response in patient-derived xenografts from gene expressions and histology images

Patient-derived xenografts (PDXs) are an appealing platform for preclinical drug studies. A primary challenge in modeling drug response prediction (DRP) with PDXs and neural networks (NNs) is the limited number of drug response samples. We investigate multimodal neural network (MM-Net) and data augmentation for DRP in PDXs. The MM-Net learns to predict response using drug descriptors, gene expressions (GE), and histology whole-slide images (WSIs). We explore whether combining WSIs with GE improves predictions as compared with models that use GE alone. We propose two data augmentation methods which allow us training multimodal and unimodal NNs without changing architectures with a single larger dataset: 1) combine single-drug and drug-pair treatments by homogenizing drug representations, and 2) augment drug-pairs which doubles the sample size of all drug-pair samples. Unimodal NNs which use GE are compared to assess the contribution of data augmentation. The NN that uses the original and the augmented drug-pair treatments as well as single-drug treatments outperforms NNs that ignore either the augmented drug-pairs or the single-drug treatments. In assessing the multimodal learning based on the MCC metric, MM-Net outperforms all the baselines. Our results show that data augmentation and integration of histology images with GE can improve prediction performance of drug response in PDXs.

60 APPLIED LIFE SCIENCES↗

A Multiagent Deep Reinforcement Learning-Enabled Dual-Branch Damping Controller for Multimode Oscillation

Here, this study develops a multiagent deep reinforcement learning (MADRL)-enabled framework for the decentralized cooperative control of a novel dual-branch (DB) damping controller for both low-frequency oscillation (LFO) and ultralow-frequency oscillation (ULFO). It has two branches, each of which consists of a proportional resonance (PR) and a second-order polynomial that is designed to handle target oscillation modes. To improve the robustness of the controller to system uncertainties, MADRL is developed, where multiagents are centrally trained to obtain the coordinated adaptive control policy while being executed in a decentralized manner to provide the optimal parameter setting for each controller with only local states. Comparisons with the IEEE 10-machine 39-bus system demonstrate that the proposed method achieves better robustness to uncertainties, lower communication delay, and single-point failure, as well as damping control performances for both LFO and ULFO.

97 MATHEMATICS AND COMPUTING↗

Exploring Causal Physical Mechanisms via Non-Gaussian Linear Models and Deep Kernel Learning: Applications for Ferroelectric Domain Structures

Rapid emergence of multimodal imaging in scanning probe, electron, and optical microscopies has brought forth the challenge of understanding the information contained in these complex data sets, targeting the intrinsic correlations between different channels, and further exploring the underpinning causal physical mechanisms. Here, we develop such an analysis framework for Piezoresponse Force Microscopy. We argue that under certain conditions, we can bootstrap experimental observations with the prior knowledge of materials structure to get information on certain nonobserved properties, and demonstrate linear causal analysis for PFM observables. We further demonstrate that the strength of individual causal links between complex descriptors can be ascertained using the deep kernel learning (DKL) model. In this DKL analysis, we use the prior information on domain structure within the image to predict the physical properties. This analysis demonstrates the correlative relationships between morphology, piezoresponse, elastic property, etc., at nanoscale. The prediction of morphology and other physical parameters illustrates a mutual interaction between surface condition and physical properties in ferroelectric materials. Overall, this analysis is universal and can be extended to explore the correlative relationships of other multichannel data sets, and allow for high-fidelity reconstruction of underpinning functionalities and physical mechanisms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗