Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Vision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

CRADA Final Report: CRADA Number NFE-24-10036 with ThermaMatrix, Inc.

ThermaMatrix, Inc provides novel vision inspection solutions for a wide range of manufacturers and industries, providing and implementing the leading technologies for nondestructive inspection (NDI) and material characterization. Many other inspection solutions are either not adequate or are not approachable due to implementation barriers needing expert level operators, excessive inspection time, and high cost. ThermaMatrix’s advanced vision inspection technology addresses all of these limitations. The Lab Embedded Entrepreneurial Program (LEEP) opportunity by the Department of Energy (DOE) allows small-business start-ups to leverage national laboratory capabilities and skilled scientists to rapidly develop their technology that aligns with DOE goals. ThermaMatrix, Inc. was positioned in the Innovation Crossroads program at Oak Ridge National Laboratory to further develop the novel Watson Vision Inspection System to support manufacturing quality control efforts. The research goals were (1) explore fundamental parameters that would improve preexisting capabilities, (2) full-scale industrial setup for demonstration, and (3) capability testing and verification. Manufacturing is demanding more NDI implementation to support their quality control needs, which this technology development would support.

36 MATERIALS SCIENCE↗

Automated Fire Detection for Industrial Settings with Pretrained Convolutional Networks

Early fire detection in industrial environments is critical to preventing equipment damage, personal injury, and operational disruptions. Traditional smoke detectors, while effective, often experience delays due to the time required for smoke to reach sensors, allowing fires to spread. Manual fire watch operations and human surveillance of camera feeds are resource-intensive and prone to human error. To address these challenges, this paper explores the application of convolutional neural networks for automated fire detection, specifically in industrial settings. By leveraging 11 different pre-trained machine vision models from TensorFlow and enhancing them with transfer learning on a custom-built industrial fire dataset, we optimized fire detection performance. Here, we analyzed each machine vision model architecture in terms of its depth, width, and input image resolution, considering both resource requirements and detection accuracy. We further explored the option of combining multiple models into an ensemble classifier to evaluate whether the performance improvements could justify the much greater computational complexity and other practical impacts. A cost-benefit analysis is presented to evaluate the trade-offs between performance and computational expense. Our findings identify that EfficientNetV2L, specifically tailored for industrial applications, provides the optimal balance between costs involved in training and using the model versus the overall fire detection performance. Additionally, we present a qualitative analysis of model performance using the technique of gradient-based class activation mapping to provide explainability by visualizing model decisions.

artificial intelligence↗

Advancements in reflected target nonintrusive assessment (ReTNA) for large optical surface measurement

Reflected computer vision targets are a powerful tool for measurement of mirror surface shape, with several important advantages over traditional fringe deflectometry methods. This method was first presented in 2021 and has undergone significant improvement and demonstration since. We describe a new baseline system using reflected computer vision targets, and present results from a large-scale measurement campaign conducted on both commercial heliostats and test mirrors in the laboratory. Calibration of the measurement system with photogrammetry allows for accurate measurement without careful control of target shape or camera position. Overall, the results show that a baseline setup using this method achieves measurement uncertainties in the slope error root-mean-square less than ±0.11 milliradian due to a series of repeatability conditions, varying sample position, rotation, lighting, camera settings, and system rebuild and recalibration. We present a detailed description of the setup, the results generated by this measurement tool, repeated measurement results, and the strengths and limitations of this metrology system.

14 SOLAR ENERGY↗

Block segmentation in feature space for realtime object detection in high granularity images

Computer vision has applications in object detection, image recognition and classification, and object tracking. One of the challenges of computer vision is the presence of useful information at multiple distance scales. Filtering techniques may sacrifice details at small scales in order to prioritize the analysis of large-scale features of the image. We present a strategy for coarse-graining multidimensional data while maintaining fine-grained detail for subsequent analysis. The algorithm is based on fixed-size block segmentation in the feature space. We apply this strategy to solve the long-standing challenge of detecting particle trajectories at the Large Hadron Collider in real time.

Computer vision↗

Micrometer: Micromechanics transformer for predicting full field mechanical responses of heterogeneous materials

Predicting mechanical responses of heterogeneous materials across scales remains a significant challenge. Traditional computational methods often struggle with complex and multiscale nature of these materials, limiting their effectiveness in real-world applications. Here, in this paper, we introduce Micrometer, a vision transformer based deep learning model designed to predict full field mechanical responses of heterogeneous materials, bridging the gap between computer vision and solid mechanics problems. We show that Micrometer, trained on a large-scale high-resolution dataset of 2D fiber-reinforced composites, can achieve state-of-the-art performance in predicting microscale strain fields across a wide range of material properties and loading conditions. Our model demonstrates accuracy and computational efficiency in applications such as computational homogenization and multiscale modeling, reducing computational time by up to two orders of magnitude compared to conventional numerical solvers while maintaining less than 1 % errors in predicting macroscale stress fields. Furthermore, we showcase Micrometer’s adaptability through transfer learning experiments on new materials with limited data, highlighting its potential to tackle diverse scenarios in computational solid mechanics. These results represent a significant step towards AI-driven innovation in materials science, addressing the limitations of traditional numerical methods and paving the way for more efficient simulations of heterogeneous materials across various industrial applications.

Composite materials↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗

Anti-distortion bioinspired camera with an inhomogeneous photo-pixel array

The bioinspired camera, comprising a single lens and a curved image sensor—a photodiode array on a curved surface—, was born of flexible electronics. Its economical build lends itself well to space-constrained machine vision applications. The curved sensor, much akin to the retina, helps image focusing, but the curvature also creates a problem of image distortion, which can undermine machine vision tasks such as object recognition. Here we report an anti-distortion single-lens camera, where 4096 silicon photodiodes arrayed on a curved surface in a nonuniform pattern assimilated to the distorting optics are the key to anti-distortion engineering. That is, the photo-pixel distribution pattern itself is warped in the same manner as images are warped, which correctively reverses distortion. Acquired images feature no appreciable distortion across a 120° horizontal view, as confirmed by their neural-network recognition accuracies. This distortion correction via photo-pixel array reconfiguration is a form of in-sensor computing.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Fourier-based three-dimensional multistage transformer for aberration correction in multicellular specimens

High-resolution tissue imaging is often compromised by sample-induced optical aberrations that degrade resolution and contrast. Although wavefront sensor-based adaptive optics (AO) can measure these aberrations, such hardware solutions are typically complex, expensive to implement and slow when serially mapping spatially varying aberrations across large fields of view. Here we introduce AOViFT (adaptive optical vision Fourier transformer)—a machine learning-based aberration sensing framework built around a three-dimensional multistage vision transformer that operates on Fourier domain embeddings. AOViFT infers aberrations and restores diffraction-limited performance in puncta-labeled specimens with substantially reduced computational cost, training time and memory footprint compared to conventional architectures or real-space networks. We validated AOViFT on live gene-edited zebrafish embryos, demonstrating its ability to correct spatially varying aberrations using either a deformable mirror or postacquisition deconvolution. By eliminating the need for the guide star and wavefront sensing hardware and simplifying the experimental workflow, AOViFT lowers technical barriers for high-resolution volumetric microscopy across diverse biological samples.

Alshaabi, Thayer [Howard Hughes Medical Institute,↗

DUNE Phase II: scientific opportunities, detector concepts, technological solutions

The international collaboration designing and constructing the Deep Underground Neutrino Experiment (DUNE) at the Long-Baseline Neutrino Facility (LBNF) has developed a two-phase strategy toward the implementation of this leading-edge, large-scale science project. The 2023 report of the US Particle Physics Project Prioritization Panel (P5) reaffirmed this vision and strongly endorsed DUNE Phase I and Phase II, as did the European Strategy for Particle Physics. While the construction of the DUNE Phase I is well underway, this White Paper focuses on DUNE Phase II planning. DUNE Phase-II consists of a third and fourth far detector (FD) module, an upgraded near detector complex, and an enhanced 2.1 MW beam. The fourth FD module is conceived as a “Module of Opportunity”, aimed at expanding the physics opportunities, in addition to supporting the core DUNE science program, with more advanced technologies. This document highlights the increased science opportunities offered by the DUNE Phase II near and far detectors, including long-baseline neutrino oscillation physics, neutrino astrophysics, and physics beyond the standard model. It describes the DUNE Phase II near and far detector technologies and detector design concepts that are currently under consideration. A summary of key R&D goals and prototyping phases needed to realize the Phase II detector technical designs is also provided. DUNE's Phase II detectors, along with the increased beam power, will complete the full scope of DUNE, enabling a multi-decadal program of groundbreaking science with neutrinos.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

Introducing SpaceNet 9 - Cross-Modal Satellite Imagery Registration for Natural Disaster Responses

Computer vision algorithms are increasingly leveraged to accelerate geospatial analysis for disaster response and recovery. As the diversity of remote sensing imagery grows with optical, SAR, and other modalities, a perquisite for analytics is cross-modal image registration. There is a high potential to harness computer vision for this pre-processing requirement toward enabling downstream analytics such as heterogeneous change detection, automated feature extraction, and data fusion. Advancement in these areas has the potential to simplify data wrangling tasks and further accelerate disaster response timelines. The SpaceNet 9 challenge (launching in mid-2024) focuses on addressing the cross-modal image registration problem and demonstrating the utility of such modules on earthquake impacted scenarios. This paper describes the motivation for the SpaceNet 9 and provides a first overview of the dataset, the baseline algorithm, and implications for seeking cross-modal image registration in Earth observation. Code is available at https://github.com/SpaceNetChallenge/SpaceNet9.

Hansch, Ronny↗

Improving Robustness of Spectrogram Classifiers with Neural Stochastic Differential Equations

Signal analysis and classification is fraught with high levels of noise and perturbation. Computer-vision-based deep learning models applied to spectrograms have proven useful in the field of signal classification and detection; however, these methods aren't designed to handle the low signal-to-noise ratios inherent within non-vision signal processing tasks. While they are powerful, they are currently not the method of choice in the inherently noisy and dynamic critical infrastructure domain, such as smart-grid sensing, anomaly detection, and non-intrusive load monitoring. Currently, these models can be brittle, which makes them susceptible to noisy input. This also means they have sub-optimal stability of explanation outputs. Experts and technicians using these models to make decisions in real world scenarios need assurance that a model is performing as it is supposed to. The classification or prediction outputs it generates should be sound and grounded, not likely to change in the presence of shifting noise landscapes. In this work, we explore the idea of Neural Stochastic Differential Equations (NSDE's) to improve the robustness of models trained to classify time series data and the effect of NSDE's on the explainability of outputs. We then test the effectiveness of these approaches by applying them to a non-intrusive load monitoring (NILM) dataset that consists of simulated harmonic signals injected into a real building.

Brogan, Joel↗

Location generalizability of image-based air quality models

This paper is to be submitted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Computer Vision for Earth Observation workshop. The full paper abstract is below: The ability to rapidly quantify atmospheric pollutants is important both for global emissions monitoring and for mitigating the adverse effects that follow a hazardous chemical release. In the aftermath of a chemical release, imagery is often the only available resource to assess local conditions. Recent work has demonstrated initial success in predicting particulate matter pollution from imagery; however, these results are tied to a specific site and do not generalize to new geographic locations. In this work, we seek to understand how easily deep learning models generalize to new locations in the context of image-based air quality assessments, targeting two distinct tasks: (1) broad measures of particulate matter pollution, and (2) the mass of a given chemical released in hazardous plumes. For the latter, we focus on sulfur dioxide, a toxic aerosol and a major component of particulate matter pollution caused by industrial fossil fuel consumption. To develop a model that operates in the widest possible range of environments, we test different training strategies, including the use of new geolocation foundation models. The best performing models achieve >80% accuracy when evaluating unseen imagery at previously seen sites, but we find significant drops in performance when evaluating imagery from unseen sites, at best 65%. Additionally, we present the public release of the National Parks Air Quality Index Dataset, a new medium-sized dataset that pairs imagery with sensor-based air quality measurements at 15 different national parks.

Byler, Eleanor B. [BATTELLE (PACIFIC NW LAB)]↗

YOLO11 to SAM2 pipeline for feature extraction from nuclear test films

The response to the effects of nuclear detonations is supported by models that describe the evolution of the nuclear fireball and cloud and the associated transport of active debris. Validation of those descriptions relies on data from the nuclear test operations. Video records of those events offer a rich source of information that was exploited to a limited extent in historic analyses. Computer vision and machine learning techniques are powerful tools that can be used to increase the number of measurements that can be obtained from those films. In this work, we apply computer vision techniques to automatically track the temporal evolution of the nuclear fireball. In particular, we apply You Only Look Once 11 (YOLO11) and Segment Anything Model 2 (SAM2) in combination with minimal human intervention to digitized versions of the original nuclear test films. As part of the proposed workflow, the YOLO11 model is applied to films to determine bounding boxes for the fireball within each frame. These are then used as inputs to SAM2, which uses image segmentation to determine the fireball boundaries and their temporal evolution. We assess the accuracy of our approach by using it to determine the energy released during the Trinity nuclear test and comparing the results with previous analyses based on manual measurements.

Van Exel, Kimberly [ORNL] (ORCID:0009000877463894)↗

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277↗

Visual Instance-aware Prompt Tuning

Visual Prompt Tuning (VPT) has emerged as a parameter-efficient fine-tuning paradigm for vision transformers, with conventional approaches utilizing dataset-level prompts that remain the same across all input instances. We observe that this strategy results in sub-optimal performance due to high variance in downstream datasets. To address this challenge, we propose Visual Instance-aware Prompt Tuning (ViaPT), which generates instance-aware prompts based on each individual input and fuses them with dataset-level prompts, leveraging Principal Component Analysis (PCA) to retain important prompting information. Moreover, we reveal that VPT-Deep and VPT-Shallow represent two corner cases based on a conceptual understanding, in which they fail to effectively capture instance-specific information, while random dimension reduction on prompts only yields performance between the two extremes. Instead, ViaPT overcomes these limitations by balancing dataset-level and instance-level knowledge, while reducing the amount of learnable parameters compared to VPT-Deep. Extensive experiments across 34 diverse datasets demonstrate that our method consistently outperforms state-of-the-art baselines, establishing a new paradigm for analyzing and optimizing visual prompts for vision transformers.

Xiao, Xi [ORNL] (ORCID:0009000009316982)↗

Double Visual Defense

This is the official code for the paper "Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness". This code can be used to produce vision language models (VLMs), like LLaVA, with enhanced robustness to adversarial attacks (e.g. jailbreaks).

Bartoldson, Brian [Lawrence Livermore National Lab↗