Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multimodal data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study

# 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge Study The 2024 University of Puerto Rico at Mayagüez Civic Innovation Challenge (CIVIC) Study provided insight into the travel patterns and associated energy consumption of participants. Study results helped researchers identify opportunities for the development of shared mobility strategies—such as collaborative ride-sharing programs—that could address the mobility needs of rural communities in Puerto Rico. The Civic Innovation Challenge is a multiagency, federal government research and action competition that funds ready-to-implement, research-based pilot projects that have the potential for scalable, sustainable, and transferable impact on community-identified priorities. ## Data Collection Agency The University of Puerto Rico at Mayagüez conducted the study. ## Survey Methodology The study was conducted in Spanish. Data collection was enabled via the open-source [NREL OpenPATH platform](https://www.nrel.gov/transportation/openpath). The resulting dataset consists of partially automated travel diaries—combining sensed and surveyed data reflecting patterns of multimodal, end-to-end, individual human mobility—as well as demographic and socioeconomic information from the 31 participants. ## Survey Records, Data, and Documentation Survey records include 31 participants. The total number of trips was 1,373 and the total non-air-miles traveled was approximately 8,260.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Exploring Causal Physical Mechanisms via Non-Gaussian Linear Models and Deep Kernel Learning: Applications for Ferroelectric Domain Structures

Rapid emergence of multimodal imaging in scanning probe, electron, and optical microscopies has brought forth the challenge of understanding the information contained in these complex data sets, targeting the intrinsic correlations between different channels, and further exploring the underpinning causal physical mechanisms. Here, we develop such an analysis framework for Piezoresponse Force Microscopy. We argue that under certain conditions, we can bootstrap experimental observations with the prior knowledge of materials structure to get information on certain nonobserved properties, and demonstrate linear causal analysis for PFM observables. We further demonstrate that the strength of individual causal links between complex descriptors can be ascertained using the deep kernel learning (DKL) model. In this DKL analysis, we use the prior information on domain structure within the image to predict the physical properties. This analysis demonstrates the correlative relationships between morphology, piezoresponse, elastic property, etc., at nanoscale. The prediction of morphology and other physical parameters illustrates a mutual interaction between surface condition and physical properties in ferroelectric materials. Overall, this analysis is universal and can be extended to explore the correlative relationships of other multichannel data sets, and allow for high-fidelity reconstruction of underpinning functionalities and physical mechanisms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Real-time Object Bounding in LiDAR Data With Computer Vision

The Multimodal Measurement System is a roadside radiation measurement testbed used to detect radiation sources in passing vehicles. It works by combining sensor signals from various modalities to produce a thorough scan of the source. A LiDAR sensor is used to measure the dimensions of the vehicle and provide a velocity estimate. However, the current LiDAR setup uses propriety software for which the source code is unavailable and cannot be updated to improve performance. Therefore, it is imperative to the accuracy of the analysis to create a custom vehicle detection that can return the dimensions and velocity of passing vehicles in real time. This new custom detection is written in C++ using the PointCloud Library, which keeps it lightweight. It also utilizes Docker and the Robot Operating System, which allows the versatility of running both on a small computer or the Lawrence Livermore National Laboratory cluster while utilizing different models of LiDAR sensors. The custom detection outperforms the current detection model, which increases the accuracy of radiation source detection.

97 MATHEMATICS AND COMPUTING↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

Alternative Vegetation States in Tropical Forests and Savannas: The Search for Consistent Signals in Diverse Remote Sensing Data

Globally, the spatial distribution of vegetation is governed primarily by climatological factors (rainfall and temperature, seasonality, and inter-annual variability). The local distribution of vegetation, however, depends on local edaphic conditions (soils and topography) and disturbances (fire, herbivory, and anthropogenic activities). Abrupt spatial or temporal changes in vegetation distribution can occur if there are positive (i.e., amplifying) feedbacks favoring certain vegetation states under otherwise similar climatic and edaphic conditions. Previous studies in the tropical savannas of Africa and other continents using the MODerate Resolution Imaging Spectroradiometer (MODIS) vegetation continuous fields (VCF) satellite data product have focused on discontinuities in the distribution of tree cover at different rainfall levels, with bimodal distributions (e.g., concentrations of high and low tree cover) interpreted as alternative vegetation states. Such observed bimodalities over large spatial extents may not be evidence for alternate states, as they may include regions that have different edaphic conditions and disturbance histories. In this study, we conduct a systematic multi-scale analysis of diverse MODIS data streams to quantify the presence and spatial consistency of alternative vegetation states in Sub-Saharan Africa. The analysis is based on the premise that major discontinuities in vegetation structure should also manifest as consistent spatial patterns in a range of remote sensing data streams, including, for example, albedo and land surface temperature (LST). Our results confirm previous observations of bimodal and multimodal distributions of estimated tree cover in the MODIS VCF. However, strong disagreements in the location of multimodality between VCF and other data streams were observed at 1 km scale. Results suggest that the observed distribution of VCF over vast spatial extents are multimodal, not because of local-scale feedbacks and emergent bifurcations (the definition of alternative states), but likely because of other factors including regional scale differences in woody dynamics associated with edaphic, disturbance, and/or anthropogenic processes. These results suggest the need for more in-depth consideration of bifurcation mechanisms and thus the likely spatial and temporal scales at which alternative states driven by different positive feedback processes should manifest.

Savanna↗

Automated electrosynthesis reaction mining with multimodal large language models (MLLMs)

Leveraging the chemical data available in legacy formats such as publications and patents is a significant challenge for the community. Automated reaction mining offers a promising solution to unleash this knowledge into a learnable digital form and therefore help expedite materials and reaction discovery. However, existing reaction mining toolkits are limited to single input modalities (text or images) and cannot effectively integrate heterogeneous data that is scattered across text, tables, and figures. In this work, we go beyond single input modalities and explore multimodal large language models (MLLMs) for the analysis of diverse data inputs for automated electrosynthesis reaction mining. We compiled a test dataset of 65 articles (MERMES-T24 set) and employed it to benchmark five prominent MLLMs against two critical tasks: (i) reaction diagram parsing and (ii) resolving cross-modality data interdependencies. The frontrunner MLLM achieved ≥96% accuracy in both tasks, with the strategic integration of single-shot visual prompts and image pre-processing techniques. We integrate this capability into a toolkit named MERMES (multimodal reaction mining pipeline for electrosynthesis). Our toolkit functions as an end-to-end MLLM-powered pipeline that integrates article retrieval, information extraction and multimodal analysis for streamlining and automating knowledge extraction. This work lays the groundwork for the increased utilization of MLLMs to accelerate the digitization of chemistry knowledge for data-driven research.

Leong, Shi Xuan↗

An Agenda for Multimodal Foundation Models for Earth Observation

Archives of remote sensing (RS) data are increasing swiftly as new sensing modalities with enhanced spatiotemporal resolution become operational. While promising new breakthroughs, the sheer volume of RS archives stretches the limits of human analysts and existing AI tools, as most models are: i) limited to single data modalities; ii) task-specific; iii) heavily reliant on labeled data. The emerging Foundation Models (FMs) have the potential to address these limitations. Trained on vast unlabeled datasets through self-supervised learning, FMs enable generic feature extraction that facilitate specialization to a wide variety of downstream tasks. This paper describes a vision towards an FM for multimodal Earth Observation data (FM4EO), discussing key building blocks and open challenges. We put particular emphasis on multimodal reasoning, a topic underexplored in EO. Our ultimate goal is a practical path toward FM4EO with capacity to unlock breakthroughs in few-shot learning scenarios, multimodal geographic knowledge integration, synthesis, and hypothesis generation.

Ambrozio Dias, Philipe↗

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES↗

Robotics for HVAC applications: A critical review and future perspectives

Recent advances in artificial intelligence (AI), enhanced computational capabilities, and innovations in sensors and hardware have driven the increasing development and application of robots in heating, ventilation, and air conditioning (HVAC) systems. We selected and reviewed 101 studies published between 2005 and 2025, sourced from IEEE Xplore, Scopus, Web of Science, and the ACM Digital Library. To analyze these works, we developed a five-dimensional analytical framework (morphology, sensing, navigation, task execution, and system integration), inspired by the Springer Handbook of Robotics and tailored specifically for robotic applications in HVAC. Based on the reviewed studies, six distinct tasks spanning the entire HVAC lifecycle have been identified. Among the six tasks, inspection and maintenance dominate (59 %), followed by indoor monitoring and auditing (21 %), whereas leakage detection, comfort support, and installation/retrofit remain less explored. To address the identified gaps, this review proposes future research directions including investigating robot-aware HVAC design principles, developing multimodal HVAC sensing and data fusion techniques, enhancing robot training and hardware capabilities, and expanding robotic applications beyond Maintenance and Operations (M&O). The findings from this review inform future robotics research for HVAC applications and ultimately enhance system affordability, energy efficiency, resilience or reliability, and occupant environmental comfort. Moreover, it seeks to inspire researchers to explore the intersections of robotics, computer science, building science, and HVAC engineering fostering advancements in this multidisciplinary field.

AI↗

In situ spin coater for multimodal grazing incidence x-ray scattering studies

We present herein a custom-made, in situ, multimodal spin coater system with an integrated heating stage that can be programmed with spinning and heating recipes and that is coupled with synchrotron-based, grazing-incidence wide- and small-angle x-ray scattering. The spin coating system features an adaptable experimental chamber, with the ability to house multiple ancillary probes such as photoluminescence and visible optical cameras, to allow for true multimodal characterization and correlated data analysis. This system enables monitoring of structural evolutions such as perovskite crystallization and polymer self-assembly across a broad length scale (2 Å–150 nm) with millisecond temporal resolution throughout a complete thin film fabrication process. The use of this spin coating system allows scientists to gain a deeper understanding of temporal processes of a material system, to develop ideal conditions for thin film manufacturing.

47 OTHER INSTRUMENTATION↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Application of the cubic polynomial strength criterion to the failure analysis of composite materials

A comparative failure analysis is presented based on the application of quadratic and cubic forms of the tensor polynomial lamina strength criterion to various composite structural configurations in a plane stress state. Failure loads have been predicted for off-angle laminates under simple loading conditions and for symmetric-balanced laminates subject to varying degrees of biaxial tension, including configurations subject to multimode failures. Some experimental data are also provided to support these calculations. From these results, the necessity of employing a cubic strength criterion to accurately predict the failure of composite laminae is demonstrated.

Tennyson, R. C.↗

Discriminative analysis of schizophrenia patients using graph convolutional networks: A combined multimodal MRI and connectomics analysis

Introduction Recent studies in human brain connectomics with multimodal magnetic resonance imaging (MRI) data have widely reported abnormalities in brain structure, function and connectivity associated with schizophrenia (SZ). However, most previous discriminative studies of SZ patients were based on MRI features of brain regions, ignoring the complex relationships within brain networks. Methods We applied a graph convolutional network (GCN) to discriminating SZ patients using the features of brain region and connectivity derived from a combined multimodal MRI and connectomics analysis. Structural magnetic resonance imaging (sMRI) and resting-state functional magnetic resonance imaging (rs-fMRI) data were acquired from 140 SZ patients and 205 normal controls. Eighteen types of brain graphs were constructed for each subject using 3 types of node features, 3 types of edge features, and 2 brain atlases. We investigated the performance of 18 brain graphs and used the TopK pooling layers to highlight salient brain regions (nodes in the graph). Results The GCN model, which used functional connectivity as edge features and multimodal features (sMRI + fMRI) of brain regions as node features, obtained the highest average accuracy of 95.8%, and outperformed other existing classification studies in SZ patients. In the explainability analysis, we reported that the top 10 salient brain regions, predominantly distributed in the prefrontal and occipital cortices, were mainly involved in the systems of emotion and visual processing. Discussion Our findings demonstrated that GCN with a combined multimodal MRI and connectomics analysis can effectively improve the classification of SZ at an individual level, indicating a promising direction for the diagnosis of SZ patients. The code is available at https://github.com/CXY-scut/GCN-SZ.git .

Chen, Xiaoyi↗

A multimodal large language model for materials science

Understanding and predicting the properties of inorganic materials is crucial for accelerating advancements in materials science and driving applications in energy, electronics and beyond. Integrating material structure data with language-based information through multimodal large language models (LLMs) offers great potential to support these efforts by enhancing human–artificial intelligence interaction. However, a key challenge lies in integrating atomic structures at full resolution into LLMs. In this work, we introduce MatterChat, a versatile structure-aware multimodal LLM that unifies material structural data and textual inputs into a single cohesive model. MatterChat uses a bridging module to effectively align a pretrained universal machine learning interatomic potential with a pretrained LLM, reducing training costs and enhancing flexibility. Our results demonstrate that MatterChat greatly improves performance in material property prediction and human–artificial intelligence interaction, surpassing general-purpose LLMs such as GPT-4. We also demonstrate its usefulness in applications such as more advanced scientific reasoning and step-by-step material synthesis.

Tang, Yingheng [Lawrence Berkeley National Laborat↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗