Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multimodal deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

CrossMP: Enabling Cross-Modality Translation between Single-Cell RNA-Seq and Single-Cell ATAC-Seq through Web-Based Portal

In recent years, there has been a growing interest in profiling multiomic modalities within individual cells simultaneously. One such example is integrating combined single-cell RNA sequencing (scRNA-seq) data and single-cell transposase-accessible chromatin sequencing (scATAC-seq) data. Integrated analysis of diverse modalities has helped researchers make more accurate predictions and gain a more comprehensive understanding than with single-modality analysis. However, generating such multimodal data is technically challenging and expensive, leading to limited availability of single-cell co-assay data. Here, we propose a model for cross-modal prediction between the transcriptome and chromatin profiles in single cells. Our model is based on a deep neural network architecture that learns the latent representations from the source modality and then predicts the target modality. It demonstrates reliable performance in accurately translating between these modalities across multiple paired human scATAC-seq and scRNA-seq datasets. Additionally, we developed CrossMP, a web-based portal allowing researchers to upload their single-cell modality data through an interactive web interface and predict the other type of modality data, using high-performance computing resources plugged at the backend.

59 BASIC BIOLOGICAL SCIENCES↗

Resolution-enhanced X-ray fluorescence microscopy via deep residual networks

Abstract Multimodal hard X-ray scanning probe microscopy has been extensively used to study functional materials providing multiple contrast mechanisms. For instance, combining ptychography with X-ray fluorescence (XRF) microscopy reveals structural and chemical properties simultaneously. While ptychography can achieve diffraction-limited spatial resolution, the resolution of XRF is limited by the X-ray probe size. Here, we develop a machine learning (ML) model to overcome this problem by decoupling the impact of the X-ray probe from the XRF signal. The enhanced spatial resolution was observed for both simulated and experimental XRF data, showing superior performance over the state-of-the-art scanning XRF method with different nano-sized X-ray probes. Enhanced spatial resolutions were also observed for the accompanying XRF tomography reconstructions. Using this probe profile deconvolution with the proposed ML solution to enhance the spatial resolution of XRF microscopy will be broadly applicable across both functional materials and biological imaging with XRF and other related application areas.

36 MATERIALS SCIENCE↗

Autonomous continuous flow reactor synthesis for scalable atom-precision

With new instrumentation design, robotics, and in-operando hyphenated analytical tool automation, the intelligent discovery of synthesis pathways is becoming feasible. It can potentially bridge the gap for the scale-up of new materials. In this article, we review current progress and describe a new system that uses an autonomous continuous flow chemistry framework to translate high-quality lead molecules and materials to quantities that can meet scalability demands. At the core is a continuous flow synthesis platform that can design its viable synthesis pathway to a particular molecule or material and then autonomously carry it out. This is realized by integrating: (1) A workflow/architecture for multimode chemical/materials characterization in-line. The in-line characterization modes are NMR, ESR, IR, Raman, UV-Vis, GC-MS, and HPLC, along with ex-situ modes for X-Ray and neutron scattering; (2) Integration for feedback/analysis/data storage of the control variables; (3) A core software stack that includes deep learning and reinforcement learning alongside quantum chemistry and molecular dynamics; (4) On-demand compute architectures that parse calculations to compute resources needed which include light-weight edge, mid-level edge (NVIDA DGX-2), and high-performance computing. We demonstrate preliminary results on how this autonomous reactor system can enhance our ability to deliver deuterated materials, copolymers, and site-substituted molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating template generation in resonant anomaly detection searches with optimal transport

We introduce Resonant Anomaly Detection with Optimal Transport (RAD-OT), a method for generating signal templates in resonant anomaly detection searches. RAD-OT leverages the fact that the samples from the conditional probability density of the target features vary approximately linearly along the optimal transport path connecting the resonant feature. This does not assume that the conditional density itself is linear with the resonant feature, allowing RAD-OT to efficiently capture multimodal relationships, changes in resolution, etc. By solving the optimal transport problem, RAD-OT can quickly build a template by interpolating between the background distributions in two sideband regions. We demonstrate the performance of RAD-OT using the LHC Olympics R&D dataset, where we find comparable sensitivity and improved stability with respect to deep learning-based approaches.

Automation↗

Postearthquake Damage Mapping via Remote Sensing: Lessons From the 2023 Türkiye Disaster

This review addresses the urgent need for scalable, accurate, and reproducible remote sensing solutions following the February 2023 Türkiye earthquakes. It synthesizes the contributions of five peer-reviewed studies published in the IEEE JSTARS Special Issue on postearthquake damage and risk assessment. These studies cover areas such as damage classification with deep learning, fusion of multisource remote sensing data, creation of benchmark datasets, detailed damage mapping, and analysis of geophysical signals using outgoing longwave radiation. The article summarizes the methodological approaches and the practical relevance of the reviewed studies for detecting, evaluating, and quantifying damage, and outlines key challenges, including model generalization, class ambiguity, and data integration. It also discusses emerging trends, including explainable artificial intelligence, multimodal data fusion, and open-data platforms. This synthesis provides a foundation for building robust, interpretable, and real-time disaster response systems and aims to guide future research in earthquake-related Earth observation and rapid damage assessment.

Taskin, Gulsen [Istanbul Technical University] (OR↗

A Multifidelity and Multimodal Machine Learning Approach for Extracting Bonding Environments of Impurities and Dopants from X-ray Spectroscopies

Extended X-ray absorption fine structure (EXAFS) spectroscopy is crucial for determining the coordination environment of impurities and dopants; however, it requires difficult measurements. X-ray absorption near edge structure (XANES) spectroscopy and X-ray emission spectroscopy (XES) can be obtained easily but cannot be converted to determine structures. In this work we develop tools to map measured XANES to the EXAFS signal through machine learning, thereby facilitating the use of EXAFS structural-determination analyses on XANES data. Through the use of Deep Operator Networks (DeepONets), we are able to accurately predict the EXAFS spectrum between 6 and 14 Å -1 from the first 6 Å -1 (~100 eV) of the absorption spectrum of Cu 2+ substitutional defects in the Fe 3+ mineral hematite (a-Fe 2 O 3 ). This surprising finding implies that theoretical analyses of X-ray absorption spectra could be implemented that extract the same conclusions as high-quality EXAFS studies from spectra collected over a much smaller range of photon energies. To encourage similar efforts, the simulated x-ray spectra, machine learning, and fitting code is made publicly available.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Avoiding fusion plasma tearing instability with deep reinforcement learning

For stable and efficient fusion energy production using a tokamak reactor, it is essential to maintain a high-pressure hydrogenic plasma without plasma disruption. Therefore, it is necessary to actively control the tokamak based on the observed plasma state, to manoeuvre high-pressure plasma while avoiding tearing instability, the leading cause of disruptions. This presents an obstacle-avoidance problem for which artificial intelligence based on reinforcement learning has recently shown remarkable performance. However, the obstacle here, the tearing instability, is difficult to forecast and is highly prone to terminating plasma operations, especially in the ITER baseline scenario. Previously, we developed a multimodal dynamic model that estimates the likelihood of future tearing instability based on signals from multiple diagnostics and actuators. Here we harness this dynamic model as a training environment for reinforcement-learning artificial intelligence, facilitating automated instability prevention. We demonstrate artificial intelligence control to lower the possibility of disruptive tearing instabilities in DIII-D, the largest magnetic fusion facility in the United States. The controller maintained the tearing likelihood under a given threshold, even under relatively unfavourable conditions of low safety factor and low torque. In particular, it allowed the plasma to actively track the stable path within the time-varying operational space while maintaining H-mode performance, which was challenging with traditional preprogrammed control. This controller paves the path to developing stable high-performance operational scenarios for future use in ITER.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A robust and interpretable machine learning approach using multimodal biological data to predict future pathological tau accumulation

The early stages of Alzheimer’s disease (AD) involve interactions between multiple pathophysiological processes. Although these processes are well studied, we still lack robust tools to predict individualised trajectories of disease progression. Here, we employ a robust and interpretable machine learning approach to combine multimodal biological data and predict future pathological tau accumulation. In particular, we use machine learning to quantify interactions between key pathological markers (β-amyloid, medial temporal lobe atrophy, tau and APOE 4) at mildly impaired and asymptomatic stages of AD. Using baseline non-tau markers we derive a prognostic index that: (a) stratifies patients based on future pathological tau accumulation, (b) predicts individualised regional future rate of tau accumulation, and (c) translates predictions from deep phenotyping patient cohorts to cognitively normal individuals. Our results propose a robust approach for fine scale stratification and prognostication with translation impact for clinical trial design targeting the earliest stages of AD.

60 APPLIED LIFE SCIENCES↗

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie↗

DCTRGAN: improving the precision of generative models with reweighting

Significant advances in deep learning have led to more widely used and precise neural network-based generative models such as Generative Adversarial Networks (Gans). We introduce a post-hoc correction to deep generative models to further improve their fidelity, based on the Deep neural networks using the Classification for Tuning and Reweighting (Dctr) protocol. The correction takes the form of a reweighting function that can be applied to generated examples when making predictions from the simulation. We illustrate this approach using Gans trained on standard multimodal probability densities as well as calorimeter simulations from high energy physics. We show that the weighted Gan examples significantly improve the accuracy of the generated samples without a large loss in statistical power. This approach could be applied to any generative model and is a promising refinement method for high energy physics applications and beyond.

47 OTHER INSTRUMENTATION↗

Multimodal imaging and machine learning to enhance microscope images of shale

A machine learning based image processing workflow is presented to enhance shale source rock microscopic images obtained using diverse imaging platforms. Images were acquired from a 30 μm diameter cylindrical Vaca Muerta shale sample using both nondestructive Transmission X-Ray Microscopy (TXM, alternately referred to as nano computed tomography) and destructive Focused Ion Beam-Scanning Electron Microscopy (FIB-SEM). Output cross-sectional images from each modality were aligned using a combination of manual and automated registration techniques to create a registered image dataset. We then apply this dataset for two image processing tasks: prediction of image cross sections with SEM-like resolution from nondestructive TXM data and repair of charged region artifacts (localized accumulation of electrons) within SEM images. The image processing algorithms for both tasks use deep learning models, specifically image-to-image Convolutional Neural Networks (CNNs) and conditional Generative Adversarial Networks (cGANs). In the image enhancement tasks, we are able to achieve significant qualitative and quantitative improvement in TXM images. Here, the best model reaches an average Peak Signal to Noise Ratio (PSNR) of 15.8 dB. Conditioning on TXM data is also shown to reduce artifacts from SEM charging, achieving an average PSNR of 25.8 dB. Furthermore, our results suggest that properly trained and validated networks are capable of significant enhancement of images obtained using nondestructive techniques, thereby improving interpretation of two- and three-dimensional images while preserving samples for future use.

58 GEOSCIENCES↗

Machine Learning for Joint Quality Control

The use of lightweight material combinations has been highly demanded in manufacturing automotive structures. However, making robust dissimilar material joints of such lightweight materials is still challenging. A significant barrier to achieving high-quality and repeatable joint performance is a deficient understanding of the relationship between the welding process, joint attributes, and joint performance. In this context, welding factors refer to material, equipment, environment, and process parameters, while joint features comprise specific microstructural attributes of the weld such as nugget size, heat affected zone (HAZ) topology, intermetallic layer thickness, and sheet thickness reduction. Joint performance is quantified in terms of strength (e.g., tensile shear, coach peel, cross-tension), weld size, and hardness, among other factors. While there have been many attempts to establish this process-structure-property relationship by developing a model derived from the associated physics and first principles, the complexity of the joining processes compounded by the complex interactions with different materials in an automotive assembly line environment, has hindered the usefulness of such attempts. The complexity is further exacerbated using different stacking materials, especially comprising dissimilar material combinations. In practice, the common approach has been the laborious process of creating welds, characterizing them, and then physically testing them through experimentation. With the emergence of artificial intelligence (AI) methods, an alternative pathway to eliciting the desired process-structure-property relationship at an accelerated pace is to use a data-driven approach by employing machine-learning (ML) techniques. This approach is benefitted by the availability of large streams of data, generated through years of research and testing by original equipment manufacturers, in the form of material, process, environmental, equipment, microstructural, and bulk-scale performance information from multimodal, multiscale sensors making measurements from laboratory-scale to production-scale processes. During Phase I efforts, which ended in fiscal year (FY) 2021, the Oak Ridge National Laboratory and Pacific Northwest National Laboratory (ORNL/PNNL) team demonstrated the effectiveness of different ML/AI frameworks in modeling complex relationships between resistance spot welding (RSW) process parameters, weld attributes, and joint properties using a subset of data from General Motors (GM). In FY 2022, the project team further refined and expanded their respective ML models to analyze additional welds with new weld stack-ups and materials to enhance the ML model predictive capability. ORNL extended its unified deep neural networks (DNN) ML training and prediction framework with new data streams of process parameters, and PNNL extended its model describing RSW process parameters’ associations with weld attributes. In FY 2023, the project team completed the development of the AI/ML architecture for analyzing aluminum/steel joints manufactured by GM via RSW and transitioned into the inline welding quality monitoring task for steel/steel RSW joints provided by GM.

36 MATERIALS SCIENCE↗

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN↗

Evaluating Deception Detection Model Robustness To Linguistic Variation

With the increasing use of automated, machine learning-driven tools and the downstream impact that algorithmic judgements can have, it is critical to develop models that are robust to evolving or manipulated inputs. Evaluating the reliability of multimodal models across linguistic variations to understand model susceptibility to intentional linguistic adversarial attacks as well as natural linguistic variations is essential in this pursuit. We present extensive analysis of model robustness and susceptibility to linguistic variations in the setting of deceptive news detection, a difficult classification task that is an increasingly important problem to solve with the impact of misinformation spread online. We evaluate the effectiveness of incorporating adversarial defense strategies and measure model susceptibility to state-of-the-art adversarial attacks using two types of linguistic attacks — character and word perturbations. We consider two multiclass prediction tasks — a 3-way classification of tweets as trustworthy, propaganda, or disinformation; and a 4-way classification as clickbait, hoax, satire, or conspiracy — and compare the performance of three embeddings that have been state-of-the-art for several NLP tasks — GloVe, ELMo, and BERT — to highlight consistent trends in susceptibility, high confidence misclassifications, and high impact failures. We find that character or mixed ensemble models are the most effective defense mechanisms and that character perturbations are a more effective attack than word perturbations for deception classification.

adversarial evaluation↗

Automated Experiments of Local Non–Linear Behavior in Ferroelectric Materials

An automated experiment in multimodal imaging to probe structural, chemical, and functional behaviors in complex materials and elucidate the dominant physical mechanisms that control device function is developed and implemented. Here, the emergence of non-linear electromechanical responses in piezoresponse force microscopy (PFM) is explored. Non-linear responses in PFM can originate from multiple mechanisms, including intrinsic material responses often controlled by domain structure, surface topography that affects the mechanical phenomena at the tip-surface junction, and the presence of surface contaminants. Using an automated experiment to probe the origins of non-linear behavior in ferroelectric lead titanate (PTO) and ferroelectric Al 0.93 B 0.07 N films, it is found that PTO shows asymmetric nonlinear behavior across a/c domain walls and a broadened high nonlinear response region around c/c domain walls. In contrast, for Al 0.93 B 0.07 N, well-poled regions show high linear piezoelectric responses, when paired with low non-linear responses regions that are multidomain show low linear responses and high nonlinear responses. It is shown that formulating dissimilar exploration strategies in deep kernel learning as alternative hypotheses allows for establishing the preponderant physical mechanisms behind the non-linear behaviors, suggesting that automated experiments can potentially discern between competing physical mechanisms. Furthermore, this technique can also be extended to electron, probe, and chemical imaging.

36 MATERIALS SCIENCE↗

Toward more-robust, AI-enabled subsurface seismic imaging for geotechnical applications

Non-invasive seismic imaging has the potential to cost-effectively evaluate large volumes of subsurface material to inform geotechnical site investigation. However, seismic imaging using full waveform inversion (FWI) requires significant computational time and is dependent on an initial starting model. As a result, FWI has not yet been widely adopted into geotechnical practice. Previous efforts, on relatively simple two-layered models, indicate that data-driven artificial intelligence (AI) models may be as effective as FWI at predicting 2D images of shear wave velocity (V s ). Furthermore, the AI model predictions can be made almost instantaneously after data acquisition and do not require an initial starting model. We examine the generality of these findings by developing a new AI model for subsurface seismic imaging, whereby we make several notable contributions. First, we architect a multimodal AI model that combines time- and frequency-domain representations of the seismic wavefield to predict a 50 m by 20 m subsurface image of V s . Second, we developed a new diverse dataset of 100,000 images with their corresponding seismic wavefields to train the AI model. Third, we propose four physics-informed data augmentations for data-driven seismic imaging. Fourth, we develop two prediction consistency tests to evaluate the model’s performance when the true subsurface is unknown. Our final model, which has been made publicly available, is capable of predicting a subsurface V s image from a single seismic wavefield with an average, mean absolute percent error (MAPE) of 24 %. The predictive model is applied to a field dataset and shown to be consistent with local geology and shear-wave refraction measurements from the same location.

Artificial intelligence↗

Neural-Network-Enhanced COTSIM: Advancing Predictive Capabilities for Fast DIII-D Simulations

Sustaining fusion reactions in tokamaks requires heating plasma to thermonuclear temperatures while maintaining confinement and stability. Neutral beam injection (NBI) provides heating, current drive, torque, and fueling, while electron cyclotron (EC) waves are widely used for heating and current drive; together, these actuators shape the plasma current, temperature, and density profiles. The control-oriented tokamak simulator (COTSIM), a predictive, control-oriented code, has been enhanced with neural-network surrogates for transport and sources. Turbulent transport is predicted by MMMnet—a neural-network version of the updated multimode model (MMM 9.0.10)—with significantly reduced computation time relative to MMM; neoclassical transport follows the Chang–Hinton model. NUBEAMnet, a surrogate of the Monte Carlo NUBEAM module, predicts beam-driven heating, current, and torque. EC heating and current drive use a control-oriented, empirically scaled source model; plasma resistivity follows the Spitzer formulation; bootstrap current uses the Sauter model. Equilibrium is computed using both prescribed and fixed-boundary solvers (FBSs), and the pedestal structure is modeled with an empirical pedestal model. For a representative DIII-D discharge, COTSIM predicts electron and ion temperature and safety-factor profiles in close agreement with TRANSP predictive and interpretive simulations while extending predictions through the pedestal region to the plasma edge (versus 80% of the minor radius in TRANSP). Furthermore, the equivalent COTSIM simulation runs in under 3 min compared to about 2 h for TRANSP, enabling rapid scenario planning, optimization of tokamak operation, and between-pulse control design.

Control-oriented tokamak simulator (COTSIM)↗

Simulating Atmospheric Processes in Earth System Models and Quantifying Uncertainties With Deep Learning Multi‐Member and Stochastic Parameterizations

Abstract Deep learning is a powerful tool to represent subgrid processes in climate models, but many application cases have so far used idealized settings and deterministic approaches. Here, we develop stochastic parameterizations with calibrated uncertainty quantification to learn subgrid convective and turbulent processes and surface radiative fluxes of a superparameterization embedded in an Earth System Model (ESM). We explore three methods to construct stochastic parameterizations: (a) a single Deep Neural Network (DNN) with Monte Carlo Dropout; (b) a multi‐member parameterization; and (c) a Variational Encoder Decoder with latent space perturbation. We show that the multi‐member parameterization improves the representation of convective processes, especially in the planetary boundary layer, compared to individual DNNs. The respective uncertainty quantification illustrates that methods (b) and (c) are advantageous compared to a dropout‐based DNN parameterization regarding the spread of convective processes. Hybrid simulations with our best‐performing multi‐member parameterizations remained challenging and crash within the first days. Therefore, we develop a pragmatic partial coupling strategy relying on the superparameterization for condensate emulation. Partial coupling reduces the computational efficiency of hybrid Earth‐like simulations but enables model stability over 5 months with our multi‐member parameterizations. However, our hybrid simulations exhibit biases in thermodynamic fields and differences in precipitation patterns. Despite this, the multi‐member parameterizations enable improvements in reproducing tropical extreme precipitation compared to a traditional convection parameterization. Despite these challenges, our results indicate the potential of a new generation of multi‐member machine learning parameterizations leveraging uncertainty quantification to improve the representation of stochasticity of subgrid effects.

Behrens, Gunnar [Deutsches Zentrum für Luft‐ und R↗