Engineering PapersSearch

SEARCH · Engineering Papers

Results for “multimode”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Potential Adoption and Benefits of Co-Optimized Multimode Engines and Fuels for U.S. Light-Duty Vehicles

Exploring a diverse portfolio of technologies for decarbonization is crucial to understanding the potential impacts of different technological solutions and their associated environmental implications. Using high-octane, high-sensitivity biofuel blends in co-optimized multimode engines can increase engine efficiency and reduce vehicle emissions. Here, the multimode engine research focuses on the benefits of light-duty vehicle engines, which can operate in multiple modes depending on the vehicle's load. Low-temperature combustion can improve efficiency and reduce emissions (such as those from oxides of nitrogen and particulate matter) during low-load operation, while spark ignition performance is maintained in high-load operation. These advanced engines can be optimized to run on blends of biobased fuels. This analysis models scenarios for potential market adoption of co-optimized multimode vehicles fueled by three different bioblendstocks: ethanol, isopropanol, and isobutanol. An integrated modeling approach is used to forecast the energy and environmental impacts of the deployment of co-optimized multimode vehicles and fuels in the light-duty sector over the 2020-to-2050 time horizon. The multidisciplinary approach combines vehicle sales modeling, system dynamics modeling of the biorefining industry, and life cycle assessment to estimate the emissions and energy benefits. The models consider market forces such as consumer preferences for vehicle attributes, biofuel supply and demand dynamics subject to biorefinery capacity build-out and bioresource constraints, and forecasted changes to the U.S. bulk energy system over time. Market adoption of co-optimized vehicles is evaluated across a wide parameter space for incremental vehicle cost and engine efficiency improvement. This analysis reveals that the deployment of co-optimized multimode fuels and vehicles results in up to a 5% reduction in annual sector-wide life cycle greenhouse gas (GHG) emissions by 2050, relative to a business-as-usual scenario, but is also indicates environmental trade-offs, such as higher life cycle water-use. Emission benefits could potentially increase beyond 2050, as the new technologies penetrate the market and gain a foothold. Results also show that, under certain circumstances, vehicles with engines co-optimized for use with high-octane, high-sensitivity biofuel blends can be cost-competitive with conventional gasoline, while reducing GHG emissions. Our modeling results indicate that co-optimized multimode fuels and engines can be strategically leveraged in tandem with electrification to decarbonize the light-duty sector. Co-optimized vehicles could play a role in the early years of the time horizon, while electric vehicles (EVs) could become more competitive in the later years, highlighting the complementary benefits of these technologies for GHG reductions.

Oke, Doris

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES

Fast Sideband Control of a Multimode Cavity Memory with Weak Dispersive Coupling to a Transmon

Mitigating ancilla-mediated error channels is a critical challenge in controlling high-quality superconducting cavities using circuit quantum electrodynamics (cQED). We address this by weakening the dispersive coupling while demonstrating fast, high-fidelity multimode control through transmon-mediated sideband interactions. We implement transmon-cavity SWAP gates with speeds up to 30 times larger than the bare dispersive coupling. Combined with transmon rotations, this enables universal state preparation in a single mode, though achieving unitary gates and extending control to multiple modes remains a challenge. In this work, we overcome this limitation by introducing two control strategies: (i) a shelving technique that stores populations in sideband-transparent states, and (ii) a method that exploits the dispersive shift to implement photon-number-selective transmon-cavity SWAP gates. We use these protocols to prepare Fock and binomial code states across any of the ten modes of a multimode cavity with millisecond coherence times—serving as a multimode quantum memory. We demonstrate unitaries that encode and decode an arbitrary qubit state from the transmon into corresponding vacuum and Fock state superpositions, as well as entangled NOON states of cavity mode pairs—a scheme extendable to arbitrary multimode Fock encodings in the cavity modes. Furthermore, we implement a new binomial encoding gate that converts arbitrary transmon superpositions into binomial code states in any cavity mode at a rate exceeding the dispersive shifts in our system, achieving an average post-selected state fidelity of 96.3% in a 4 μ⁢s gate time. By using precalibrated transmon and sideband pulses, our work demonstrates multimode control with significantly reduced calibration overhead, enabling efficient unitary operations using sideband interactions in multimode cQED systems.

Huang, Jordan [Rutgers Univ., Piscataway, NJ (Unit

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)

Joint Optimization of Multimodal Transit Frequency and Shared Autonomous Vehicle Fleet Size with Hybrid Metaheuristic and Nonlinear Programming

Shared autonomous vehicles (SAVs) bring competition to traditional transit services but redesigning multimodal transit network can utilize SAVs as feeders to enhance service efficiency and coverage. This paper presents an optimization framework for the joint multimodal transit frequency and SAV fleet size problem, a variant of the transit network frequency setting problem. The objective is to maximize total transit ridership (including SAV-fed trips and subtracting boarding rejections) across multiple time periods under budget constraints, considering endogenous mode choice (transit, point-to-point SAVs, driving) and route selection, while allowing for strategic route removal by setting frequencies to zero. Due to the problem’s non-linear, non-convex nature and the computational challenges of large-scale networks, we develop a hybrid solution approach that combines a metaheuristic approach (particle swarm optimization) with nonlinear programming for local solution refinement. To ensure computational tractability, the framework integrates analytical approximation models for SAV waiting times based on fleet utilization, multimodal network assignment for route choice, and multinomial logit mode choice behavior, bypassing the need for computationally intensive simulations within the main optimization loop. Applied to the Chicago metropolitan area’s multimodal network, our method illustrates a 33.3% increase in transit ridership through optimized transit route frequencies and SAV integration, particularly enhancing off-peak service accessibility and strategically reallocating resources.

Ng, Max

FIRM: federated image reconstruction using multimodal tomographic data

Here, we propose a federated algorithm for reconstructing images using multimodal tomographic data sourced from dispersed locations, addressing the challenges of traditional unimodal approaches that are prone to noise and reduced image quality, as well as the limitations of centralized multimodal approaches that require extensive data transfer, leading to significant communication overhead, storage demands, and potential data privacy concerns. Our approach formulates a joint inverse optimization problem incorporating multimodality constraints and solves it in a federated framework through local gradient computations complemented by lightweight central operations, thereby ensuring data decentralization. Leveraging the connection between our federated algorithm and the quadratic penalty method, we introduce an adaptive step-size rule with guaranteed sublinear convergence. Numerical results demonstrate superior computational efficiency and improved image reconstruction quality compared to existing approaches.

federated algorithm

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES

Multimodality in the Search for New Physics in Pulsar Timing Data and the Case of Kination-amplified Gravitational-wave Background from Inflation

We investigate the kination-amplified inflationary gravitational-wave background (GWB) interpretation of the signal recently reported by various pulsar timing array (PTA) experiments. Kination is a post-inflationary phase in the expansion history dominated by the kinetic energy of some scalar field, characterized by a stiff equation of state w = 1. Within the inflationary GWB model, we identify two modes that can fit the current data sets (NANOGrav and EPTA) with equal likelihood: the kination-amplification (KA) mode and the ordinary, no-kination-amplification (no-KA) mode. The multimodality of the likelihood motivates a Bayesian analysis with nested sampling. We analyze the free spectra of current PTA data and mock free spectra constructed with higher signal-to-noise ratios using nested sampling. The analysis of the mock spectrum designed to be consistent with the best fit to the NANOGrav 15 yr (NG15) data successfully reveals the expected bimodal posterior for the first time while excluding the reheating mode that appears in the fit to the current NG15 data, making a case for our correct and comprehensive treatment of potential multimodal posteriors arising from future PTA data sets. The resultant Bayes factor is $\mathcal{B}$ $\equiv$ Z no–KA /Z KA = 2.9 ± 1.9, indicating comparable statistical significance between the two modes. Given the theoretical model-building challenges of producing highly blue-tilted primordial tensor spectra, the KA mode has the advantage of requiring less blue primordial spectra, compared with the no-KA mode. The synergy between future cosmic microwave background polarization, pulsar timing, and laser interferometer measurements of gravitational waves will help resolve the ambiguity implied by the multimodal posterior in PTA-only searches.

Cosmology

SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving the LLMs to align existing foundation models with scientific disciplines, concepts and goals. In this work, we present SciTune as a tuning framework to improve the ability of LLMs to follow scientific multimodal instructions. To test our methodology, we use a human-generated scientific instruction tuning dataset and train a large multimodal model LLaMA-SciTune that connects a vision encoder and LLM for science-focused visual and language understanding. LLaMA-SciTune significantly outperforms the state-of-the-art models in the generated figure types and captions in multiple scientific multimodal benchmarks. In comparison to the models that are fine-tuned with machine generated data only, LLaMA-SciTune surpasses human performance on average and in many sub-categories on the ScienceQA benchmark.

• Artificial intelligence (AI) / machine learning

Fast Sideband Control of a Weakly Coupled Multimode Bosonic Memory

Circuit quantum electrodynamics (cQED) with superconducting cavities coupled to nonlinear circuits like transmons offers a promising platform for hardware-efficient quantum information processing. We address critical challenges in realizing this architecture by weakening the dispersive coupling while also demonstrating fast, high-fidelity multimode control by dynamically amplifying gate speeds through transmon-mediated sideband interactions. This approach enables transmon-cavity SWAP gates, for which we achieve speeds up to 30 times larger than the bare dispersive coupling. Combined with transmon rotations, this allows for efficient, universal state preparation in a single cavity mode, though achieving unitary gates and extending control to multiple modes remains a challenge. In this work, we overcome this by introducing two sideband control strategies: (1) a shelving technique that prevents unwanted transitions by temporarily storing populations in sideband-transparent transmon states and (2) a method that exploits the dispersive shift to synchronize sideband transition rates across chosen photon-number pairs to implement transmon-cavity SWAP gates that are selective on photon number. We leverage these protocols to prepare Fock and binomial code states across any of ten modes of a multimode cavity with millisecond cavity coherence times. We demonstrate the encoding of a qubit from a transmon into arbitrary vacuum and Fock state superpositions, as well as entangled NOON states of cavity mode pairs— a scheme extendable to arbitrary multimode Fock encodings. Furthermore, we implement a new binomial encoding gate that converts arbitrary transmon superpositions into binomial code states in $\qty{4}{\micro\second}$ (less than $1/\chi$), achieving an average post-selected final state fidelity of $\qty{96.3}{\percent}$ across different fiducial input states.

Huang, Jordan [Rutgers U., Piscataway]

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)

Cascading economic losses from port disruptions under capacity constrained multimodal freight networks

This study quantifies how throughput disruptions at major seaports cascade through capacity-constrained multimodal freight networks and interregional production systems. We couple an agent-based model (ABM) multimodal freight simulation that resolves rerouting, terminal queueing, and inventory drawdown under binding modal and facility capacities with a multiregional output loss input-output (MRIIM) model that propagates realized delivery shortfalls across regions and sectors. The framework is demonstrated for the Port of Los Angeles using Freight Analysis Framework flows and Bureau of Economic Analysis input-output accounts and is evaluated over a 52-week horizon under deterministic sector targeted shocks and stochastic disruption realizations with uncertain severity and duration. Results indicate nonlinear amplification: realized national losses concentrate in manufacturing and transportation/warehousing even when exogenous port shocks are dispersed, suggesting that congestion spillback and limited short-run substitution can dominate the initial shock allocation. We further evaluate a tabular reinforcement-learning (Q-learning) intervention layer that selects among a small set of implementable system level levers (truck-to-rail and truck-to-barge shift settings) without overriding shipper routing, finding that such interventions reduce total losses for moderate disruptions but yield diminishing returns once substitute modes approach capacity. By linking operational freight behavior to system wide impacts under uncertainty, the proposed ABM-MRIIM pipeline provides a reusable workflow for port disruption stress testing, identification of structurally critical sectors/corridors, and evaluation of resilience interventions under realistic capacity limits.

42 ENGINEERING

Revealing Local Structures through Machine-Learning-Fused Multimodal Spectroscopy

Atomistic structures of materials offer valuable insights into their functionality. Determining these structures remains a fundamental challenge in materials science, especially for systems with defects. While both experimental and computational methods exist, each has limitations in resolving nanoscale structures. Core-level spectroscopies, such as X-ray absorption (XAS) or electron energy-loss spectroscopies (EELS), have been used to determine the local bonding environment and structure of materials. Recently, machine learning (ML) methods have been applied to extract structural and bonding information from XAS/EELS data. However, frameworks relying solely on a single data stream, defined as characterization data derived from a single element using one technique, are often insufficient because multiple local environments can yield similar spectral features, making it challenging to differentiate between competing structural hypotheses. Here, in this work, we address this challenge by integrating multimodal ab initio simulations, experimental data acquisition, and ML techniques for structure characterization. Our goal is to determine local structures and properties using EELS and XAS data from multiple elements and edges. To showcase our approach, we use various lithium nickel manganese cobalt (NMC) oxide compounds which are used for lithium ion batteries, including those with oxygen vacancies and antisite defects, as the sample material system. We successfully inferred local element content, ranging from lithium to transition metals, with quantitative agreement with experimental data. Beyond local element inference, we find that ML model based on multimodal spectroscopic data is able to determine whether local defects such as oxygen vacancy and antisites are present, a task which is impossible for single mode spectra or other experimental techniques. Furthermore, our framework is able to provide physical interpretability, bridging spectroscopy with the local atomic and electronic structures.

battery

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING

Pixel-Registered Multimodal Synchrotron XRF and FTIR Microscopies Reveal Salinity Stress Response Mechanisms in Pistachio

Background: Salinity is a major abiotic stress that negatively affects nearly all plant species at all stages of growth. Drought and poor-quality irrigation cause high soil salinity and salt accumulation via evaporation, reducing crop productivity. Despite its critical importance, the spatial localization of salt ions and associated biochemical changes within plants experiencing high salinity remains largely unknown. In this study, we developed a multimodal imaging pipeline to understand the impact of salinity on the pistachio rootstock UCB-1 (Pistacia atlantica x Pistacia integerrima). We directly link biochemical fingerprints in stem tissue architecture with salt ion localization to provide insights into the strategies pistachio uses to tolerate salinity. Results: We observed that Pistacia spp. exposed to high salt conditions accumulated Ca, Si, Cl, Al and Mg as hotspots within the pith, compared to the control (of which only Ca and Al co-locate). In contrast, there was a decrease in K between the control and salinity treatment. Hotspots of amide I and II were present in the cortex and pith of the salinity treated sample. Additionally, the salinity treatment resulted in an increased abundance of pectin and carbohydrates within the pith compared to the control, and the abundance of esters/carboxylic acid was greater in the salinity treatment. Conclusions: We determined that Cl and K, S and P, and biochemical components polysaccharide and pectin, esters and carboxylic acid, amide I and cellulose are the strongest drivers of salinity- treatment induced variability. In the cortex and phloem/xylem, a negative K-Ca correlation decreases in the salinity treatment. Several hotspots of elements and amide I (proteins) appear under salinity treatment, particularly in the cortex, suggesting an increase in the production of stress-related proteins (in response to high Cl) and/or structural proteins (i.e. Ca). Together, these results indicate that pistachio responds to salinity through ion compartmentalization coupled with a targeted biochemical adjustment, rather than a broadscale tissue-wide response. Overall, these novel, spatially resolved pixel-registered multimodal imaging data provide an enabling platform to understand the mechanisms of salinity tolerance in Pistacia spp and can be broadly applied to studying stress-related phenotype response in various plant tissues.

FTIR spectromicroscopy

Toward Intelligent Multimodal Holography for Real-Time Chemical Imaging of Dynamic Ion Separation

Molecular-level visualization of ion transport and separation dynamics in complex environments is crucial for advancing energy systems, water purification, and critical materials recovery. Achieving this requires imaging platforms that combine structural sensitivity, chemical specificity, and real-time operation. Digital off-axis holography (DOAH) provides high-throughput, label-free quantitative phase imaging but inherently lacks chemical selectivity. Integrating DOAH with complementary spectroscopic channels such as fluorescence or hyperspectral imaging introduces the needed molecular specificity, while also creating challenges in multimodal data fusion, synchronization, and computational throughput. Artificial intelligence offers a powerful route to address these limitations by uniting physics-based reconstruction with data-driven interpretation. In this Perspective, we outline a framework for intelligent multimodal holography and demonstrate its potential using a preliminary AI-driven test case. Raw DOAH holograms of lanthanide solutions subjected to magnetic field gradients were analyzed using multi-agent AI workflows that autonomously selected reconstruction tools, extracted NMF components, and generated scientific claims consistent with true paramagnetic and diamagnetic behavior. This demonstration shows how AI-enabled reasoning can deliver real-time chemical–structural interpretation directly from raw holograms. Together, these advances define a path toward adaptive, intelligent holography platforms capable of supporting in situ chemical separations, dynamic ion transport analysis, and next-generation interfacial science.

Ricchiuti, Giovanna