Engineering PapersSearch

SEARCH · Engineering Papers

Results for “object detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection

Object detection in remote sensing demands extensive, high-quality annotations—a process that is both labor-intensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing early-stage localization accuracy. Extensive evaluations on multiple remote sensing datasets plus a real-world user study, demonstrate that our framework not only reduces annotation effort, but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface will be made available.

Burges, Marvin [ORNL] (ORCID:0000000312690769)

Interactive Rotated Object Detection for Novel Class Detection in Remotely Sensed Imagery

In this paper we propose IRTR-DETR an Interactive and Real-Time Rotated DEtection TRansformer that extends IRTDETR to predict rotated bounding boxes. IRTR-DETR maintains the Human-In-The-Loop (HIL) workflow of IRTDETR but introduces rotation-aware heads for improved detection of objects with arbitrary orientations. Similarly to IRTDETR IRTR-DETR can be trained with a small labeled sample set in an interactive setting but we show that it can also be pretrained on related but not identical data--such as a building damage dataset--before being applied to tasks like identifying buildings under construction. We demonstrate the efficacy of our approach on the publicly available Tiny-DOTA and xBD dataset as well as two study-cases on proprietary datasets of greenhouses and houses under construction ("waffle homes"). Detecting greenhouses is highly relevant in the context of damage assessment while "waffle homes" aid understanding typical floorplans and building codes in different areas both thereby supporting population modeling emergency response and policy planning. Our method outperforms the state of the art in interactive rotated object detection on the Tiny-DOTA dataset by 5.7 percent and improves upon the non interactive RTDETR by 7.85 to 19.39 percent (depending on the number of provided samples) while maintaining its real-time efficiency.

Burges, Marvin [ORNL] (ORCID:0000000312690769)

Block segmentation in feature space for realtime object detection in high granularity images

Computer vision has applications in object detection, image recognition and classification, and object tracking. One of the challenges of computer vision is the presence of useful information at multiple distance scales. Filtering techniques may sacrifice details at small scales in order to prioritize the analysis of large-scale features of the image. We present a strategy for coarse-graining multidimensional data while maintaining fine-grained detail for subsequent analysis. The algorithm is based on fixed-size block segmentation in the feature space. We apply this strategy to solve the long-standing challenge of detecting particle trajectories at the Large Hadron Collider in real time.

Computer vision

Towards automated and real-time multi-object detection of anguilliform fishes from sonar data using YOLOv8 deep learning algorithm

Eels (Anguilla spp.), including American eels (Anguilla rostrata), European eels (Anguilla anguilla), and Japanese eels (Anguilla japonica), are species of critical management and regulatory concern due to their vulnerability to various stressors during downstream migrations. Accurate and efficient detection of migrating eels can improve our understanding of fish behaviors and fish-hydraulic structure interactions, thereby facilitating the design, operation, and optimization of more effective downstream passage facilities from both biological and economic perspectives. However, a real-time, automated framework for detecting migrating eels in real-world applications is currently lacking. Leveraging imaging sonar as a reliable technology for fish passage monitoring, field data are acquired using imaging sonar and then converted to single sonar frames/images for subsequent analysis. In this study, a framework based on the You Only Look Once Version 8 (YOLOv8)-based convolutional neural network is proposed for multi-object detection of eels and non-eel fish using the sonar images after image subtraction and additional wavelet denoising. The results from both training and testing phases demonstrate that the framework's ability can successfully detect both eels and non-eel fish in preprocessed sonar images, achieving F1-scores and mAP@0.50 exceeding 0.84. Additionally, the incorporation of wavelet denoising during preprocessing slightly improve detection performance. Furthermore, the transferability of this framework from eel to lamprey detection is demonstrated to be feasible given the similar morphological characteristics of these two species. Overall, the proposed framework achieves accurate and efficient detection of migrating eels, providing reliable and real-time information that can help conserve vulnerable eel and eel-like populations.

Deep learning

Object detection with deep learning for rare event search in the GADGET II TPC

In the pursuit of identifying rare two-particle events within the GADGET II Time Projection Chamber (TPC), this paper presents a comprehensive approach for leveraging Convolutional Neural Networks (CNNs) and various data processing methods. To address the inherent complexities of 3D TPC track reconstructions, the data is expressed in 2D projections and 1D quantities. This approach capitalizes on the diverse data modalities of the TPC, allowing for the efficient representation of the distinct features of the 3D events, with no loss in topology uniqueness. Additionally, it leverages the computational efficiency of 2D CNNs and benefits from the extensive availability of pre-trained models. Given the scarcity of real training data for the rare events of interest, simulated events are used to train the models to detect real events. To account for potential distribution shifts when predominantly depending on simulations, significant perturbations are embedded within the simulations. This produces a broad parameter space that works to account for potential physics parameter and detector response variations and uncertainties. These parameter-varied simulations are used to train sensitive 2D CNN object detectors. When combined with 1D histogram peak detection algorithms, this multi-modal detection framework is highly adept at identifying rare, two-particle events in data taken during experiment 21072 at the Facility for Rare Isotope Beams (FRIB), demonstrating a 100% recall for events of interest. Here, we present the methods and outcomes of our investigation and discuss the potential future applications of these techniques.

Convolutional neural network

Different methods of estimating riverbed sediment grain size diverge at the basin scale

Introduction: The distribution of sediment grain size in streams and rivers is often quantified by the median grain size (D50), a key metric for understanding and predicting hydrologic and biogeochemical function of streams and rivers. Manual D50 measurements are time-consuming and ignore larger grains, while approaches to model D50 based on catchment characteristics may over-generalize and miss site-scale heterogeneity. Machine learning-enabled object detection methods like You Only Look Once (YOLO) provides an alternative that enables estimation of D50 that is faster than manual measurements and more site-specific than predictions based on catchment characteristics. Methods: To understand the potential role of object detection methods for improving understanding of D50, we compared D50 estimates made manually, predicted from catchment characteristics, and using a YOLO-enabled approach across the Yakima River Basin. Results: We found distinct differences between methods for D50 averages and variability, and relationships between D50 estimates and basin characteristics. Discussion: We discuss the advantages and limitations of object detection methods versus current methods, and explore potential future directions to combine D50 methods to better estimate spatiotemporal variation of D50, and improve incorporation into basin-scale models.

grain size distribution

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Dark Energy Survey Year 6 Results: Synthetic-source Injection Across the Full Survey Using Balrog

Synthetic source injection (SSI), the insertion of sources into pixel-level on-sky images, is a powerful method for characterizing object detection and measurement in wide-field, astronomical imaging surveys. Within the Dark Energy Survey (DES), SSI plays a critical role in characterizing all necessary algorithms used in converting images to catalogs, and in deriving quantities needed for the cosmology analysis, such as object detection rates, galaxy redshift estimation, galaxy magnification, star-galaxy classification, and photometric performance. We present here a source injection catalog of 146 million injections spanning the entire 5000 deg 2 DES footprint, generated using the Balrog SSI pipeline. Through this SSI sample, we demonstrate that the DES Year 6 (Y6) image processing pipeline provides accurate estimates of the object properties, for both galaxies and stars, at the percent-level, and we highlight specific regimes where the accuracy is reduced. We then show the consistency between SSI and data catalogs, for all galaxy samples developed within the weak lensing and galaxy clustering analyses of DES Y6. The consistency between the two catalogs also extends to their correlations with survey observing properties (seeing, airmass, depth, extinction, etc.). Lastly, we highlight a number of applications of this catalog to the DES Y6 cosmology analysis, such as estimates of the redshift distribution and lens magnification. This dataset is the largest SSI catalog produced at this fidelity and will serve as a key testing ground for exploring the utility of SSI catalogs in upcoming surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS

ObstacleSense: Low-Power Neuromorphic Vision for Corridor Obstacle Awareness in Low-Level ADAS

The automotive industry’s pursuit of Level 5 autonomy is constrained by substantial perception-compute power requirements, often reaching 1, 000 + watts in full autonomy stacks. Reducing this energy burden requires rethinking perception not only at the high-end autonomy level, but also at the foundational Advanced Driver Assistance Systems (ADAS) level where low-power, safety-critical sensing can have broad impact. Neuromorphic vision provides a promising starting point: HD Dynamic Vision Sensors (DVS) can operate below 100 mW at the sensor level by reporting only asynchronous brightness changes. However, low-power sensing alone is insufficient if downstream perception reintroduces dense, energy-intensive computation. In particular, many event-driven object-detection pipelines still rely on CNN backbones, while purely spiking alternatives often trade away accuracy or ignore deployment constraints. We introduce ObstacleSense, a highly compact, CNN-free hybrid ANN–SNN framework for Level 0–1 forward-corridor obstacle awareness. Instead of performing full-scene object detection with a convolutional feature backbone, ObstacleSense targets the safety-critical question of whether the ego corridor is occupied and how far the nearest obstacle is. The architecture combines polarity-conditioned event encoding, lightweight temporal spiking dynamics, axial spatial mixing, and coarse-to-fine range estimation within a regular fixed-grid compute pattern. This design avoids the dense CNN backbone commonly used in event-based detection while maintaining a small state footprint suitable for eventual small-FPGA deployment. Before hardware mapping, we evaluate the software implementation using a model-side power proxy derived from MACs, weight and activation traffic, and spiking state updates under shared FP16 assumptions. On simulated CARLA event corpora, the deployment-oriented model achieves 0.9464 objectness F1, 0.9978 grid-level mAP, and 0.8987 m distance Mean Absolute Error at an estimated 1.92 mW proxy cost, while maintaining performance on unseen generalization test sequences.

Johnson-Scott, Zac [ORNL]

3D reconstruction and neural rendering for adversarial machine learning

While evasion attacks on computer vision systems have been widely studied, creating attacks that remain effective under significant changes in viewpoint continues to be challenging. Traditional approaches often rely on affine transformations of images, but these approaches degrade at larger perspective shifts and often produce unrealistic or ineffective perturbations. Recent methods use differentiable renderers to improve viewpoint robustness, but they typically depend on manually constructed 3D models. We introduce a semi-automated pipeline that generates physically printable and perspective-invariant adversarial patches using only a small set of 2D images. Our method integrates 3D reconstruction, neural rendering, adversarial patch optimization, and an object detection victim model into a unified workflow. We use 2D Gaussian Splatting for high fidelity mesh reconstruction and FlexPara for surface parameterization that produces texture maps suitable for patch editing. Together, these components form a fully differentiable pipeline in PyTorch3D that links texture modification to model outputs, enabling efficient optimization of patches that remain effective across many viewpoints. The complete process, from image capture to patch printing and physical evaluation, can be completed within a few hours. We demonstrate the effectiveness of the resulting patches through attacks on the YOLOv8 object detection model and discuss remaining challenges and opportunities for improving robustness and scalability.

Singhvi, Vivaan [ORNL] (ORCID:0009000586288221)

Automated Construction of Artificial Lattice Structures with Designer Electronic States

Manipulating matter with a scanning tunneling microscope (STM) enables the creation of atomically defined artificial structures that host designer quantum states. However, the time-consuming nature of the manipulation process, coupled with the sensitivity of the STM tip, constrains the exploration of diverse configurations and limits the size of the designed features. In this study, we present a reinforcement learning (RL)-based framework for creating artificial structures by spatially manipulating carbon monoxide (CO) molecules on a copper substrate by using the STM tip. The automated workflow combines molecule detection and manipulation, employing deep-learning-based object detection to locate CO molecules and linear assignment algorithms to allocate these molecules to designated target sites. We initially perform molecule maneuvering based on randomized parameter sampling for sample bias, tunneling current set point, and manipulation speed. This data set is then structured into an action trajectory used to train an RL agent. The model is subsequently deployed on the STM for real-time fine-tuning of the manipulation parameters during structure construction. Our approach incorporates path-planning protocols coupled with active drift compensation to enable atomically precise fabrication of structures with significantly reduced human input while realizing larger-scale artificial lattices with the desired electronic properties. Furthermore, using our approach, we demonstrate the automated construction of an extended artificial graphene lattice and confirm the existence of a characteristic Dirac point in its electronic structure. Further challenges regarding the RL-based structural assembly scalability are discussed.

Algorithms

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics

Developing a Deep Learning-Computer Vision Framework to Monitor Avian Interactions with Solar Energy Facility Infrastructure (Final Technical Report)

The project addressed an inability to monitor avian interactions with photovoltaic (PV) solar energy facilities necessary for understanding PV solar impacts on birds. In the project, machine-vision technology that continuously monitors avian activities at PV solar facilities was developed. The technology includes four machine-learning (ML) models, each of which accomplishes a specific task in detecting birds and classifying their activities in live or recorded videos—detecting and tracking moving objects, differentiating birds from other objects, detecting bird collisions with solar panels, and classifying non-collision bird activities around PV facilities. Major project outcomes include adoption by two of DOE SETO’s SolWEB projects, providing novel observational data on birds to promote co-location of PV solar development and habitat conservation, known as ecovoltaics.

14 SOLAR ENERGY

INR-TEM: Robust cavity detection in multifocus TEM images via implicit neural representations

When characterizing materials using transmission electron microscopy (TEM) images, detecting and quantifying small features in microstructures, such as cavities, pose significant challenges. Off-the-shelf object detection models, including YOLOv8, show considerable performance degradation, particularly when images vary in resolution and the objects of interest possess a low percentage of the total image region of interest. In this study, we introduce a novel detection pipeline that incorporates an implicit neural representation (INR)-based detection method, INR-TEM, and two-modality imaging (e.g., under-focused and over-focused images typically acquired during materials characterization) to improve object detection performance. The INR-TEM method incorporates a pixel-wise prediction principle inspired by pixel-wise centerness weighting. INR-TEM demonstrates superior robustness to resolution variability, maintaining high detection accuracy even at low image resolutions compared to YOLOv8. To leverage INR-TEM effectively in real-world two-modality characterization applications, we further integrate a two-stage motion correction pipeline designed explicitly for aligning multifocus TEM images. The alignment process, comprising keypoint (based on scale-invariant feature transform, SIFT) and intensity matching, significantly mitigates the adverse effects of perceived motion-induced image degradation during through-focal TEM imaging, directly enhancing INR-TEM’s detection capability over conventional single-focus images. Our integrated INR-TEM cavity detection framework notably improves performance across various cavity sizes, outperforming off-the-shelf YOLOv8 detections that rely on a single image modality.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO

Abstract Streambed grain sizes control river hydro‐biogeochemical (HBGC) processes and functions. However, measuring their quantities, distributions, and uncertainties is challenging due to the diversity and heterogeneity of natural streams. This work presents a photo‐driven, artificial intelligence (AI)‐enabled, and theory‐based workflow for extracting the quantities, distributions, and uncertainties of streambed grain sizes from photos. Specifically, we first trained You Only Look Once, an object detection AI, using 11,977 grain labels from 36 photos collected from nine different stream environments. We demonstrated its accuracy with a coefficient of determination of 0.98, a Nash–Sutcliffe efficiency of 0.98, and a mean absolute relative error of 6.65% in predicting the median grain size of 20 ground‐truth photos representing nine typical stream environments. The AI is then used to extract the grain size distributions and determine their characteristic grain sizes, including the 10th, 50th, 60th, and 84th percentiles, for 1,999 photos taken at 66 sites within a watershed in the Northwest US. The results indicate that the 10th, median, 60th, and 84th percentiles of the grain sizes follow log‐normal distributions, with most likely values of 2.49, 6.62, 7.68, and 10.78 cm, respectively. The average uncertainties associated with these values are 9.70%, 7.33%, 9.27%, and 11.11%, respectively. These data allow for the computation of the quantities, distributions, and uncertainties of streambed HBGC parameters, including Manning's coefficient, Darcy‐Weisbach friction factor, top layer interstitial velocity magnitude, and nitrate uptake velocity. Additionally, major sources of uncertainty in grain sizes and their impact on HBGC parameters are examined.

58 GEOSCIENCES

Utilization of Data Augmentation Techniques in Automated Inspection Systems for Defect Detection in Metals With Limited Data

Accurate identification of defects on metal surfaces is of great interest to many industry sectors, such as the automotive and aerospace industries. In contrast to conventional manual inspection techniques, recent automated inspection systems employ deep learning models trained to detect defects rapidly and precisely. The development of these models often requires a substantial image dataset to acquire adequate knowledge of defect features and enhance their predictive accuracy. When data is limited, augmentation techniques are often used to improve the precision and accuracy of defect detection systems. This study examined the prediction performance of two object detection models, namely Faster Region‐based Convolutional Neural Network (Faster R‐CNN) and You Only Look Once version 8 (YOLOv8), to identify dent defects in limited images of cast iron cylinder head surfaces. The original image set contains 46 images with 563 dents. To overcome limited data availability, common image augmentation techniques along with a copy‐paste method were applied. Results show that standard augmentation improved YOLOv8 accuracy by 8.00% and average precision (AP) by 3.00%. On the other hand, the copy‐paste technique achieved a 20.00% increase in accuracy and a 1% increase in AP with just 200 synthetic dents. Furthermore, these results provide support for using the copy‐paste augmentation strategy to enhance defect detection performance, with a limited dataset, contributing to more accurate defect identification in remanufacturing processes.

36 MATERIALS SCIENCE

Automated Detection of Qubit Structures on Quantum Chips

In this study, YOLO (You Only Look Once), a well-known object detection and image segmentation model, is used to detect qubits on a quantum chip. The model is trained by exposing it to various images of objects of interest and by tweaking various training parameters to allow the model to learn from a limited and highly specialized dataset of SEM imagery. The model was provided with numerous images of qubits and various prominent features of note located on a qubit. These detections were then integrated into a simple GUI, allowing for real-time feedback alongside SEM usage. This will allow us to create various functions enabling automated functions like chip-wide scanning, auto-focusing, and image correction.

Perjuste, Ruth [Wellesley Coll.]

Automated Detection of Qubit Structures on Quantum Chips

In this study, YOLO (You Only Look Once), a well-known object detection and image segmentation model, is used to detect qubits on a quantum chip. The model is trained by exposing it to various images of objects of interest and by setting various training parameters to allow the model to learn in unconventional conditions. We were able to provide the model with numerous images of qubits and various prominent features of note located on a qubit. We then created a simple GUI to display the model's live detections that would be integrated with the SEMs. This allowed us to create various functions such as auto-imagery of the qubits across a chip, auto-focusing, auto-contrasting, automatic staging, and directional corrections, enabling users to scan and analyze qubit surfaces in a fast and effective manner.

Perjuste, Ruth [Wellesley Coll.]