Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multimodal data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Identifying Light-Duty Vehicle Travel from Large-Scale Multimodal Wearable GPS Data with Novelty Detection Algorithms

Identifying travel mode within travel survey data sets, especially light-duty vehicle (LDV) travel, is foundational, though nontrivial, to travel behavior analysis and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. They are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the data requirements. This paper proposes an LDV travel mode detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique-one-class support vector machines (OCSVMs)-and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multimodal prediction.

47 OTHER INSTRUMENTATION↗

A Multimodal Event Catalog and Waveform Data Set That Supports Explosion Monitoring from Nevada, U.S.A.

Multimodal, curated data sets and nuisance event catalogs remain rare in the explosion monitoring community relative to curated seismic data sets. The source of this relative absence is the difficultly in deploying multimodal receivers that sense the seismic, acoustic, and other modalities from multiphysics sources. We provide such a data set in this study that delivers seismic, infrasound, and electromagnetic (magnetometer) sensor records collected over a two–week period, within 255 km of a 10 ton buried chemical explosion called DAG–4 that was located at 37.1146°, –116.0693° on 22 June 2019 21:06:19.88 UTC. This catalog includes 485 seismic, seismoacoustic, and infrasound–only events that an expert analyst manually built by reviewing waveforms from 29 seismic and infrasound sensors. Our data release includes waveforms from these 29 seismic, infrasound, and seismoacoustic stations and two magnetometer stations and their station metadata. We deliver these waveforms in NNSA KB Core CSS.w format (i4) with a corresponding wfdisc table that provides the header information. Here, we expect that this data set will provide a valuable, benchmark resource to develop signal processing algorithms and explosion monitoring methods against manual, human observations.

58 GEOSCIENCES↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

O'Hare Airport roadway traffic prediction via data fusion and Gaussian process regression

This study proposes an approach of leveraging information gathered from multiple traffic data sources at different resolutions to obtain approximate inference on the traffic distribution of Chicago's O'Hare Airport area. Specifically, it proposes the ingestion of traffic datasets at different resolutions to build spatiotemporal models for predicting the distribution of traffic volume on the road network. Due to its good adaptability and flexibility for spatiotemporal data, the Gaussian process (GP) regression was employed to provide short-term forecasts using data collected by loop detectors (sensors) and supplemented by telematics data. The GP regression is used to make predictions of the distribution of the proportion of sensor data traffic volume represented by the telematics data for each location of the sensors. Consequently, the fitted GP model can be used to determine the approximate traffic distribution for a testing location outside of the training points. Policymakers in the transportation sector can find the results of this work helpful for making informed decisions relating to current and future transportation conditions in the area.

42 ENGINEERING↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

Robust Group Subspace Recovery: A New Approach for Multi-Modality Data Fusion

Robust Subspace Recovery (RoSuRe) algorithm was recently introduced as a principled and numerically efficient algorithm that unfolds underlying Unions of Subspaces (UoS) structure, present in the data. The union of Subspaces (UoS) is capable of identifying more complex trends in data sets than simple linear models. In this work, we build on and extend RoSuRe to prospect the structure of different data modalities individually. We propose a novel multi-modal data fusion approach based on group sparsity which we refer to as Robust Group Subspace Recovery (RoGSuRe). Relying on a bi-sparsity pursuit paradigm and non-smooth optimization techniques, the introduced framework learns a new joint representation of the time series from different data modalities, respecting an underlying UoS model. We subsequently integrate the obtained structures to form a unified subspace structure. The proposed approach exploits the structural dependencies between the different modalities data to cluster the associated target objects. The resulting fusion of the unlabeled sensors’ data from experiments on audio and magnetic data has shown that our method is competitive with other state of the art subspace clustering methods. The resulting UoS structure is employed to classify newly observed data points, highlighting the abstraction capacity of the proposed method.

47 OTHER INSTRUMENTATION↗

2020 Can Do Colorado E-Bike Mini Pilot Program Study

### The Colorado Energy Office conducted a mini pilot program study as part of the Can Do Colorado initiative, providing e-bikes to 13 low-income participants. The program aimed to encourage energy-efficient transportation during the COVID-19 pandemic as transit services were reduced and people were concerned about exposure. The insights garnered from this small-scale pilot study informed the design of a full-scale, 2-year pilot in locations across Colorado. For more information about the mini pilot program, see NLR's [Preliminary Results Report](https://www.nlr.gov/docs/fy21osti/79657.pdf). Micromobility options such as e-bikes offer a solution for improving energy efficiency for short-distance trips, especially in urban areas. Pedal-assist e-bikes use an electric motor and battery to help power the bike. The motor amplifies the power behind each pedal stroke, augmenting the energy you put into the bike. #### Data Collection Agency The Colorado Energy Office conducted the survey. #### Survey Methodology Participants in the program received a Momentum LaFree E+ e-bike (Class 1) and accessories at no cost and manually submitted travel data and feedback for 3 months using the CanBikeCo App. The smartphone app, developed in partnership with NLR, used a customized version of the [NLR OpenPATH platform](https://www.nlr.gov/transportation/openpath.html). #### Survey Records and Data Survey records include a total of 13 participants. This dataset contains 3 months of end-to-end, multimodal travel data manually submitted via smartphone app by 13 low-income essential workers in the greater Denver area. The data includes distance, mode (e.g., e-bike, car, transit), trip purpose, and demographic information.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Descriptor: Infrastructure Perception and Control: Multi-Sensor Object Tracking Dataset (IPC-MSOT)

Traffic intersections are crucial and challenging nodes in transportation networks where multiple lanes of vehicles and pedestrians converge. Traffic accidents often occur at traffic intersections, including a large proportion of traffic fatalities and about one-half of all traffic injuries in the United States. Object detection data were collected in 2024 across three intersections in Colorado Springs, CO, USA, over the course of multiple days and various times to induce a heterogeneous mix of traffic conditions and behaviors. The purpose of the data collection exercises was to learn various attributes about infrastructure sensors and to build a repository of high-resolution, object-level data that can be used for research and development (e.g., to develop multisensor data fusion algorithms). The Infrastructure Perception and Control:Multi-Sensor Object tracking (IPC-MSOT) dataset was collected as part of the U.S. Department of Transportation's Strengthening Mobility and Revolutionizing Transportation (SMART) project, where the city of Colorado Springs, Colorado, and the National Renewable Energy Laboratory collaborated to collect object-level trajectory data from road users using multiple types of infrastructure sensors deployed at different intersections. This dataset allows for testing of late-stage sensor fusion algorithms and their ability to ingest multimodal sensor data, and it can be utilized by traffic engineers to design and evaluate trajectory-based signal control strategies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Solving multiphysics-based inverse problems with learned surrogates and constraints

Abstract Solving multiphysics-based inverse problems for geological carbon storage monitoring can be challenging when multimodal time-lapse data are expensive to collect and costly to simulate numerically. We overcome these challenges by combining computationally cheap learned surrogates with learned constraints. Not only does this combination lead to vastly improved inversions for the important fluid-flow property, permeability, it also provides a natural platform for inverting multimodal data including well measurements and active-source time-lapse seismic data. By adding a learned constraint, we arrive at a computationally feasible inversion approach that remains accurate. This is accomplished by including a trained deep neural network, known as a normalizing flow, which forces the model iterates to remain in-distribution, thereby safeguarding the accuracy of trained Fourier neural operators that act as surrogates for the computationally expensive multiphase flow simulations involving partial differential equation solves. By means of carefully selected experiments, centered around the problem of geological carbon storage, we demonstrate the efficacy of the proposed constrained optimization method on two different data modalities, namely time-lapse well and time-lapse seismic data. While permeability inversions from both these two modalities have their pluses and minuses, their joint inversion benefits from either, yielding valuable superior permeability inversions and CO 2 plume predictions near, and far away, from the monitoring wells.

Yin, Ziyi (ORCID:0000000250248771)↗

Pixel-Registered Multimodal Synchrotron XRF and FTIR Microscopies Reveal Salinity Stress Response Mechanisms in Pistachio

Background: Salinity is a major abiotic stress that negatively affects nearly all plant species at all stages of growth. Drought and poor-quality irrigation cause high soil salinity and salt accumulation via evaporation, reducing crop productivity. Despite its critical importance, the spatial localization of salt ions and associated biochemical changes within plants experiencing high salinity remains largely unknown. In this study, we developed a multimodal imaging pipeline to understand the impact of salinity on the pistachio rootstock UCB-1 (Pistacia atlantica x Pistacia integerrima). We directly link biochemical fingerprints in stem tissue architecture with salt ion localization to provide insights into the strategies pistachio uses to tolerate salinity. Results: We observed that Pistacia spp. exposed to high salt conditions accumulated Ca, Si, Cl, Al and Mg as hotspots within the pith, compared to the control (of which only Ca and Al co-locate). In contrast, there was a decrease in K between the control and salinity treatment. Hotspots of amide I and II were present in the cortex and pith of the salinity treated sample. Additionally, the salinity treatment resulted in an increased abundance of pectin and carbohydrates within the pith compared to the control, and the abundance of esters/carboxylic acid was greater in the salinity treatment. Conclusions: We determined that Cl and K, S and P, and biochemical components polysaccharide and pectin, esters and carboxylic acid, amide I and cellulose are the strongest drivers of salinity- treatment induced variability. In the cortex and phloem/xylem, a negative K-Ca correlation decreases in the salinity treatment. Several hotspots of elements and amide I (proteins) appear under salinity treatment, particularly in the cortex, suggesting an increase in the production of stress-related proteins (in response to high Cl) and/or structural proteins (i.e. Ca). Together, these results indicate that pistachio responds to salinity through ion compartmentalization coupled with a targeted biochemical adjustment, rather than a broadscale tissue-wide response. Overall, these novel, spatially resolved pixel-registered multimodal imaging data provide an enabling platform to understand the mechanisms of salinity tolerance in Pistacia spp and can be broadly applied to studying stress-related phenotype response in various plant tissues.

FTIR spectromicroscopy↗

Multimodal Atomic Force Microscopy for the Characterization of Metallic Particulates

This study investigates the utility of multimodal, or functional, atomic force microscopy (AFM) for the characterization of individual particles and particle ensembles. In single-particle analyses, AFM imaging modes provided orthogonal insights into mechanical and magnetic properties that are not accessible through conventional electron microscopy. Although the experiments were inherently delicate and time-intensive, with an observed sample loss rate of approximately 30%, these techniques enabled the qualitative differentiation of grain structures and magnetic domains, underscoring the challenges of experimental robustness. For particle ensembles, AFM enabled the extraction of reliable 2D and 3D particle size distributions. However, efforts to chemically differentiate particles based on mechanical property contrasts were limited by scaling effects and the qualitative nature of the data. Overall, multimodal AFM offers valuable complementary information, but its application requires careful consideration of methodological limitations, particularly for sample integrity, calibration, and scalability of mechanical and magnetic measurements. This report is organized to first address the application of AFM to single-particle analysis using various imaging modalities, followed by its potential role in ensemble-level particle screening and differentiation.

36 MATERIALS SCIENCE↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN↗

Deep learning with plasma plume image sequences for anomaly detection and prediction of growth kinetics during pulsed laser deposition

Abstract Materials synthesis platforms that are designed for autonomous experimentation are capable of collecting multimodal diagnostic data that can be utilized for feedback to optimize material properties. Pulsed laser deposition (PLD) is emerging as a viable autonomous synthesis tool, and so the need arises to develop machine learning (ML) techniques that are capable of extracting information from in situ diagnostics. Here, we demonstrate that intensified-CCD image sequences of the plasma plume generated during PLD can be used for anomaly detection and the prediction of thin film growth kinetics. We develop multi-output (2 + 1)D convolutional neural network regression models that extract deep features from plume dynamics that not only correlate with the measured chamber pressure and incident laser energy, but more importantly, predict parameters of an auto-catalytic film growth model derived from in situ laser reflectivity experiments. Our results demonstrate how ML with in situ plume diagnostics data in PLD can be utilized to maintain deposition conditions in an optimal regime. Further, the predictive capabilities of plume dynamics on the kinetics of film growth or other film properties prior to deposition provides a means for rapid pre-screening of growth conditions for the non-expert, which promises to accelerate materials optimization with PLD.

36 MATERIALS SCIENCE↗

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods↗