Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multimodal data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Robotics for HVAC applications: A critical review and future perspectives

Recent advances in artificial intelligence (AI), enhanced computational capabilities, and innovations in sensors and hardware have driven the increasing development and application of robots in heating, ventilation, and air conditioning (HVAC) systems. We selected and reviewed 101 studies published between 2005 and 2025, sourced from IEEE Xplore, Scopus, Web of Science, and the ACM Digital Library. To analyze these works, we developed a five-dimensional analytical framework (morphology, sensing, navigation, task execution, and system integration), inspired by the Springer Handbook of Robotics and tailored specifically for robotic applications in HVAC. Based on the reviewed studies, six distinct tasks spanning the entire HVAC lifecycle have been identified. Among the six tasks, inspection and maintenance dominate (59 %), followed by indoor monitoring and auditing (21 %), whereas leakage detection, comfort support, and installation/retrofit remain less explored. To address the identified gaps, this review proposes future research directions including investigating robot-aware HVAC design principles, developing multimodal HVAC sensing and data fusion techniques, enhancing robot training and hardware capabilities, and expanding robotic applications beyond Maintenance and Operations (M&O). The findings from this review inform future robotics research for HVAC applications and ultimately enhance system affordability, energy efficiency, resilience or reliability, and occupant environmental comfort. Moreover, it seeks to inspire researchers to explore the intersections of robotics, computer science, building science, and HVAC engineering fostering advancements in this multidisciplinary field.

AI↗

In situ spin coater for multimodal grazing incidence x-ray scattering studies

We present herein a custom-made, in situ, multimodal spin coater system with an integrated heating stage that can be programmed with spinning and heating recipes and that is coupled with synchrotron-based, grazing-incidence wide- and small-angle x-ray scattering. The spin coating system features an adaptable experimental chamber, with the ability to house multiple ancillary probes such as photoluminescence and visible optical cameras, to allow for true multimodal characterization and correlated data analysis. This system enables monitoring of structural evolutions such as perovskite crystallization and polymer self-assembly across a broad length scale (2 Å–150 nm) with millisecond temporal resolution throughout a complete thin film fabrication process. The use of this spin coating system allows scientists to gain a deeper understanding of temporal processes of a material system, to develop ideal conditions for thin film manufacturing.

47 OTHER INSTRUMENTATION↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Discriminative analysis of schizophrenia patients using graph convolutional networks: A combined multimodal MRI and connectomics analysis

Introduction Recent studies in human brain connectomics with multimodal magnetic resonance imaging (MRI) data have widely reported abnormalities in brain structure, function and connectivity associated with schizophrenia (SZ). However, most previous discriminative studies of SZ patients were based on MRI features of brain regions, ignoring the complex relationships within brain networks. Methods We applied a graph convolutional network (GCN) to discriminating SZ patients using the features of brain region and connectivity derived from a combined multimodal MRI and connectomics analysis. Structural magnetic resonance imaging (sMRI) and resting-state functional magnetic resonance imaging (rs-fMRI) data were acquired from 140 SZ patients and 205 normal controls. Eighteen types of brain graphs were constructed for each subject using 3 types of node features, 3 types of edge features, and 2 brain atlases. We investigated the performance of 18 brain graphs and used the TopK pooling layers to highlight salient brain regions (nodes in the graph). Results The GCN model, which used functional connectivity as edge features and multimodal features (sMRI + fMRI) of brain regions as node features, obtained the highest average accuracy of 95.8%, and outperformed other existing classification studies in SZ patients. In the explainability analysis, we reported that the top 10 salient brain regions, predominantly distributed in the prefrontal and occipital cortices, were mainly involved in the systems of emotion and visual processing. Discussion Our findings demonstrated that GCN with a combined multimodal MRI and connectomics analysis can effectively improve the classification of SZ at an individual level, indicating a promising direction for the diagnosis of SZ patients. The code is available at https://github.com/CXY-scut/GCN-SZ.git .

Chen, Xiaoyi↗

A multimodal large language model for materials science

Understanding and predicting the properties of inorganic materials is crucial for accelerating advancements in materials science and driving applications in energy, electronics and beyond. Integrating material structure data with language-based information through multimodal large language models (LLMs) offers great potential to support these efforts by enhancing human–artificial intelligence interaction. However, a key challenge lies in integrating atomic structures at full resolution into LLMs. In this work, we introduce MatterChat, a versatile structure-aware multimodal LLM that unifies material structural data and textual inputs into a single cohesive model. MatterChat uses a bridging module to effectively align a pretrained universal machine learning interatomic potential with a pretrained LLM, reducing training costs and enhancing flexibility. Our results demonstrate that MatterChat greatly improves performance in material property prediction and human–artificial intelligence interaction, surpassing general-purpose LLMs such as GPT-4. We also demonstrate its usefulness in applications such as more advanced scientific reasoning and step-by-step material synthesis.

Tang, Yingheng [Lawrence Berkeley National Laborat↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗

The POINTER Imaging baseline cohort: Associations between multimodal neuroimaging biomarkers, cardiovascular health, and cognition

Abstract INTRODUCTION The U.S. Study to Protect Brain Health Through Lifestyle Intervention to Reduce Risk (U.S. POINTER) is evaluating lifestyle interventions in older adults at risk for cognitive decline and dementia. Here we characterize the baseline data set of the POINTER Imaging ancillary study. METHODS Participants underwent health and cognitive assessments and neuroimaging with multimodal positron emission tomography (PET) (beta‐amyloid [Aβ] and tau) and magnetic resonance imaging (MRI). Framingham risk score (FRS) was used to quantify cardiovascular disease (CVD) risk. RESULTS A total of 1052 participants (31% from underrepresented ethnoracial groups) were enrolled. Compared to Aβ−, Aβ+ (29%) participants were older, had higher apolipoprotein E (APOE) ε4 carriage rate and white matter hyperintensity volume, and greater temporal tau. FRS was related to MRI measures, but not AD biomarkers. FRS and tau had independent effects on cognition. DISCUSSION In this heterogenous, at‐risk cohort, CVD risk was related to more abnormal brain structure and poorer cognition, representing a putative non‐AD (Alzheimer's disease) pathway to brain injury and cognitive decline. Highlights The U.S. Study to Protect Brain Health Through Lifestyle Intervention to Reduce Risk (U.S. POINTER) cohort is enriched for cardiovascular disease (CVD) and poor lifestyle POINTER Imaging collected multimodal neuroimaging data in this unique, at‐risk cohort Amyloid burden was related to age, apolipoprotein E (APOE) ε4 carriage, and measures of disease progression Associations between amyloid and tau, and tau and cognition, were relatively weak CVD risk and tau pathology were independently related to memory

Neurosciences & Neurology↗

From Data to Insights: A Covariate Analysis of the IARPA BRIAR Dataset for Multimodal Biometric Recognition Algorithms at Altitude and Range

This paper examines covariate effects on fused whole body biometrics performance in the IARPA BRIAR dataset, specifically focusing on UAV platforms, elevated positions, and distances up to 1000 meters. The dataset includes outdoor videos compared with indoor images and controlled gait recordings. Normalized raw fusion scores relate directly to predicted false accept rates (FAR), offering an intuitive means for interpreting model results. A linear model is developed to predict biometric algorithm scores, analyzing their performance to identify the most influential covariates on accuracy at altitude and range. Weather factors like temperature, wind speed, solar loading, and turbulence are also investigated in this analysis. The study found that resolution and camera distance best predicted accuracy and findings can guide future research and development efforts in long-range/elevated/UAV biometrics and support the creation of more reliable and robust systems for national security and other critical domains.

Bolme, David↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Sputter-Deposited Mo Thin Films: Multimodal Characterization of Structure, Surface Morphology, Density, Residual Stress, Electrical Resistivity, and Mechanical Response

Multimodal datasets of materials are rich sources of information which can be leveraged for expedited discovery of process–structure–property relationships and for designing materials with targeted structures and/or properties. For this data descriptor article, we provide a multimodal dataset of magnetron sputter-deposited molybdenum (Mo) thin films, which are used in a variety of industries including high temperature coatings, photovoltaics, and microelectronics. In this dataset we explored a process space consisting of 27 unique combinations of sputter power and Ar deposition pressure. Here, the phase, structure, surface morphology, and composition of the Mo thin films were characterized by x-ray diffraction, scanning electron microscopy, atomic force microscopy, and Rutherford backscattering spectrometry. Physical properties—namely, thickness, film stress and sheet resistance—were also measured to provide additional film characteristics and behaviors. Additionally, nanoindentation was utilized to obtain mechanical load-displacement data. The entire dataset consists of 2072 measurements including scalar values (e.g., film stress values), 2D linescans (e.g., x-ray diffractograms), and 3D imagery (e.g., atomic force microscopy images). An additional 1889 quantities, including film hardness, modulus, electrical resistivity, density, and surface roughness, were derived from the experimental datasets using traditional methods. Minimal analysis and discussion of the results are provided in this data descriptor article to limit the authors’ preconceived interpretations of the data. Overall, the data modalities are consistent with previous reports of refractory metal thin films, ensuring that a high-quality dataset was generated. The entirety of this data is committed to a public repository in the Materials Data Facility.

36 MATERIALS SCIENCE↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Zero-shot and prompt-based models have excelled at visual reasoning tasks by leveraging large-scale natural image corpora, but they often fail on sparse and domain-specific scientific image data. We introduce Zenesis, a no-code interactive computer vision platform designed to reduce data readiness bottlenecks in scientific imaging workflows. Zenesis integrates lightweight multimodal adaptation for zero-shot inference on raw scientific data, human-in-the-loop refinement, and heuristic-based temporal enhancement. We validate our approach on Focused Ion Beam Scanning Electron Microscopy (FIB-SEM) datasets of catalyst-loaded membranes. Zenesis outperforms baselines, achieving an average accuracy of 0.947, Intersection over Union (IoU) of 0.858, and Dice score of 0.923 on amorphous catalyst samples; and 0.987 accuracy, 0.857 IoU, and 0.923 Dice on crystalline samples. These results represent a significant performance gain over conventional methods such as Otsu thresholding and standalone models like the Segment Anything Model (SAM). Zenesis enables effective image segmentation in domains where annotated datasets are limited, offering a scalable solution for scientific discovery.

Mukherjee, Shubhabrata↗

pnnl/Multimodal-Few-Shot

This project combines few-shot machine learning with a multimodal framework for electron microscope data quantification.

Doty, Christina↗

Heterogeneous multiphase flow properties of volcanic rocks and implications for noble gas transport from underground nuclear explosions

Of interest to the Underground Nuclear Explosion Signatures Experiment are patterns and timing of explosion-generated noble gases that reach the land surface. The impact of potentially simultaneous flow of water and gas on noble gas transport in heterogeneous fractured rock is a current scientific knowledge gap. This article presents field and laboratory data to constrain and justify a triple continua conceptual model with multimodal multiphase fluid flow constitutive equations that represents host rock matrix, natural fractures, and induced fractures from past underground nuclear explosions (UNEs) at Aqueduct and Pahute Mesas, Nevada National Security Site, Nevada, USA. Capillary pressure from mercury intrusion and direct air–water measurements on volcanic tuff core samples exhibit extreme spatial heterogeneity (i.e., variation over multiple orders of magnitude). Petrographic observations indicate that heterogeneity derives from multimodal pore structures in ash-flow tuff components and post-depositional alteration processes. Comparisons of pre- and post-UNE samples reveal different pore size distributions that are due in part to microfractures. Capillary pressure relationships require a multimodal van Genuchten (VG) constitutive model to best fit the data. Relative permeability estimations based on unimodal VG fits to capillary pressure can be different from those based on bimodal VG fits, implying the choice of unimodal vs. bimodal fits may greatly affect flow and transport predictions of noble gas signatures. The range in measured capillary pressure and predicted relative permeability curves for a given lithology and between lithologies highlights the need for future modeling to consider spatially distributed properties.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Revealing Local Structures through Machine-Learning-Fused Multimodal Spectroscopy

Atomistic structures of materials offer valuable insights into their functionality. Determining these structures remains a fundamental challenge in materials science, especially for systems with defects. While both experimental and computational methods exist, each has limitations in resolving nanoscale structures. Core-level spectroscopies, such as X-ray absorption (XAS) or electron energy-loss spectroscopies (EELS), have been used to determine the local bonding environment and structure of materials. Recently, machine learning (ML) methods have been applied to extract structural and bonding information from XAS/EELS data. However, frameworks relying solely on a single data stream, defined as characterization data derived from a single element using one technique, are often insufficient because multiple local environments can yield similar spectral features, making it challenging to differentiate between competing structural hypotheses. Here, in this work, we address this challenge by integrating multimodal ab initio simulations, experimental data acquisition, and ML techniques for structure characterization. Our goal is to determine local structures and properties using EELS and XAS data from multiple elements and edges. To showcase our approach, we use various lithium nickel manganese cobalt (NMC) oxide compounds which are used for lithium ion batteries, including those with oxygen vacancies and antisite defects, as the sample material system. We successfully inferred local element content, ranging from lithium to transition metals, with quantitative agreement with experimental data. Beyond local element inference, we find that ML model based on multimodal spectroscopic data is able to determine whether local defects such as oxygen vacancy and antisites are present, a task which is impossible for single mode spectra or other experimental techniques. Furthermore, our framework is able to provide physical interpretability, bridging spectroscopy with the local atomic and electronic structures.

battery↗

Future trends in synchrotron science at NSLS-II

In this paper, we summarize briefly some of the future trends in synchrotron science as seen at the National Synchrotron Light Source II, a new, low emittance source recently commissioned at Brookhaven National Laboratory. We touch upon imaging techniques, the study of dynamics, the increasing use of multimodal approaches, the vital importance of data science, and finally, other enabling technologies. Each are presently undergoing a time of rapid change, driving the field of synchrotron science forward at an ever increasing pace. It is truly an exciting time and one in which Roger Cowley, to whom this journal issue is dedicated, would surely be both invigorated by, and at the heart of.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Development of Multimodal Few-Shot Analytics for Electron Micrographs

Recent advances in materials data analytics have provided new avenues for determining process-structure-property (PSP) linkages in a variety of materials. Machine learning techniques including few-shot learning have increased the efficiency of classifying microscopy images for the purposes of material characterization. Attempts at creating a multimodal approach can provide further improvements to current models and help extract more salient features from data. In this vein, raw spectrum data was taken to provide an additional modality to our current pyCHIP classifier. Modifications in segmentation also show potential in improving the accuracy of the pyCHIP classifier. Classifier output was analyzed using network graphs and unsupervised clustering algorithms such as spectral clustering to detect better segmentation methods than the current “chipping” approach. We suggest that the chip selection process can be automated in the future using a combination of these techniques to enable high-throughput analyses.

36 MATERIALS SCIENCE↗