Computer vision based rock-bolt detection in orthomosaic imagery obtained in GPS-denied environments for mining safety assessments
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.
Computer vision is regarded as one of the most complex and computationally intensive problems. An integrated vision system (IVS) is a system that uses vision algorithms from all levels of processing to perform for a high level application (e.g., object recognition). An IVS normally involves algorithms from low level, intermediate level, and high level vision. Designing parallel architectures for vision systems is of tremendous interest to researchers. Several issues are addressed in parallel architectures and parallel algorithms for integrated vision systems.
This paper is to be submitted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Computer Vision for Earth Observation workshop. The full paper abstract is below: The ability to rapidly quantify atmospheric pollutants is important both for global emissions monitoring and for mitigating the adverse effects that follow a hazardous chemical release. In the aftermath of a chemical release, imagery is often the only available resource to assess local conditions. Recent work has demonstrated initial success in predicting particulate matter pollution from imagery; however, these results are tied to a specific site and do not generalize to new geographic locations. In this work, we seek to understand how easily deep learning models generalize to new locations in the context of image-based air quality assessments, targeting two distinct tasks: (1) broad measures of particulate matter pollution, and (2) the mass of a given chemical released in hazardous plumes. For the latter, we focus on sulfur dioxide, a toxic aerosol and a major component of particulate matter pollution caused by industrial fossil fuel consumption. To develop a model that operates in the widest possible range of environments, we test different training strategies, including the use of new geolocation foundation models. The best performing models achieve >80% accuracy when evaluating unseen imagery at previously seen sites, but we find significant drops in performance when evaluating imagery from unseen sites, at best 65%. Additionally, we present the public release of the National Parks Air Quality Index Dataset, a new medium-sized dataset that pairs imagery with sensor-based air quality measurements at 15 different national parks.
The potential for using computer vision as sensory feedback for robot gas-tungsten arc welding is investigated. The basic parameters that must be controlled while directing the movement of an arc welding torch are defined. The actions of a human welder are examined to aid in determining the sensory information that would permit a robot to make reproducible high strength welds. Special constraints imposed by both robot hardware and software are considered. Several sensory modalities that would potentially improve weld quality are examined. Special emphasis is directed to the use of computer vision for controlling gas-tungsten arc welding. Vendors of available automated seam tracking arc welding systems and of computer vision systems are surveyed. An assessment is made of the state of the art and the problems that must be solved in order to apply computer vision to robot controlled arc welding on the Space Shuttle Main Engine.
This report presents the findings from market research conducted for NASA’s Aerial Aid Convergent Aeronautics Solutions (CAS) exploration project, which aims to assess the current state of the market and technological readiness for Uncrewed Aerial Systems (UAS) for medical emergency first response. The research reveals a robust and rapidly growing market for UAS, with a notable emerging sector for Drones as First Responders (DFR). Despite this growth, DFR applications are currently limited by regulatory, technical, and other challenges, which restrict their use primarily to manned remote video surveillance, and therefore are primarily employed by police units. To our knowledge, there is no evidence of UAS being utilized by medical first responders for scene assessment. Limited evidence exists for closely related applications; however, these are mostly confined to pilot programs for the delivery of medical supplies or equipment. Although there has been discussion around fully autonomous DFR applications for medical purposes such as UAS ambulances or patient transport drones, these applications are generally not yet operational in practice. The technology for full autonomy, especially in guidance and control, has seen significant advancements, and recent Federal Aviation Administration (FAA)regulations are likely to accelerate adoption. Computer vision algorithms for fully autonomous medical emergency response scene surveillance are primed for advancement and deployment. A notable gap likely exists between advancements in computer vision research and what is being integrated in the commercial DFR sector. This gap is primarily due to challenges such as quality assurance for autonomous systems, the availability of application-specific training datasets for computer vision algorithms, regulatory constraints, and public perception and privacy concerns.
High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.
In 1991, famous French scientist Pierre-Gilles de Genes was awarded Nobel prize for his impactful research in soft matter, more specifically polymers. He is defined as the founding father of soft matter. In his Nobel lecture (https://www.nobelprize.org/uploads/2018/06/gennes-lecture.pdf ) he described soft matter aka complex fluids as materials with two primary features – (a) complexity and (b) flexibility. The sub-categories of soft matter (e.g.- granular materials, polymers, foams, colloids etc.) are defined on the basis of Pierre-Gilles de Gennes’ definition. At NASA GRC, we are pushing the boundaries for fundamental study of soft matter on Lunar Surface. With regard to Lunar surface science, we are focusing on developing capabilities pertaining to granular materials and bio-soft/active matter to facilitate future efforts in ISRU and bio-ISRU capabilities. In order to achieve fundamental goals of soft matter research within the limitations of Lunar environment, the scientific capabilities need to be small, flexible, modular, off the shelf and the focus needs to be more on developing an interdisciplinary capability that leverages the recent growth in AI/ML and Computer Vision to augment our understanding of fundamental science. This strategy would allow us to reduce our resource requirement during launch, installation, and occupied real estate footprint on Lunar surface In this talk, we will go over 3 different capabilities that we have developed in house and in close collaboration – (a) Differential Dynamic Microscopy (DDM), (b) Portable In-situ Chemical Spectroscopy (PICS) and (c) Computer Vision Enabled Observation. At very high level, Differential Dynamic Microscopy (DDM) allows us to study the structure-property-process relation (microrheology) of bio-soft/active matter using optical microscope and improved image analysis capabilities. PICS uses AI/ML-based advanced signal deconvolution and analysis technique that can work with existing portable spectroscopy tools to perform materials analysis (e.g.- granular materials and bio-soft/active matter) inspection on the go. Finally, computer vision enabled analysis allows us to use simple camera images for 3D reconstruction of experimental process and tracking of objects of interest in an experiment. We expect that this detailed process will allow us reach a thorough understanding of soft matter in Lunar environment. The capabilities developed by us will help to validate and establish fundamental understanding in Lunar environment. This will, in turn, allow us to guide future space exploration missions and expand the knowledge base of the scientific and engineering communities.
Vision is examined in terms of a computational process, and the competence, structure, and control of computer vision systems are analyzed. Theoretical and experimental data on the formation of a computer vision system are discussed. Consideration is given to early vision, the recovery of intrinsic surface characteristics, higher levels of interpretation, and system integration and control. A computational visual processing model is proposed and its architecture and operation are described. Examples of state-of-the-art vision systems, which include some of the levels of representation and processing mechanisms, are presented.
Computer vision has applications in object detection, image recognition and classification, and object tracking. One of the challenges of computer vision is the presence of useful information at multiple distance scales. Filtering techniques may sacrifice details at small scales in order to prioritize the analysis of large-scale features of the image. We present a strategy for coarse-graining multidimensional data while maintaining fine-grained detail for subsequent analysis. The algorithm is based on fixed-size block segmentation in the feature space. We apply this strategy to solve the long-standing challenge of detecting particle trajectories at the Large Hadron Collider in real time.
The automation of low-altitude rotorcraft flight depends on the ability to detect, locate, and navigate around obstacles lying in the rotorcraft's intended flightpath. Computer vision techniques provide a passive method of obstacle detection and range estimation, for obstacle avoidance. Several algorithms based on computer vision methods have been developed for this purpose using laboratory data; however, further development and validation of candidate algorithms require data collected from rotorcraft flight. A data base containing low-altitude imagery augmented with the rotorcraft and sensor parameters required for passive range estimation is not readily available. Here, the emphasis is on the methodology used to develop such a data base from flight-test data consisting of imagery, rotorcraft and sensor parameters, and ground-truth range measurements. As part of the data preparation, a technique for obtaining the sensor calibration parameters is described. The data base will enable the further development of algorithms for computer vision-based obstacle detection and passive range estimation, as well as provide a benchmark for verification of range estimates against ground-truth measurements.
The advanced inspection system is an autonomous control and analysis system that improves the inspection and remediation operations for ground and surface systems. It uses optical imaging technology with intelligent computer vision algorithms to analyze physical features of the real-world environment to make decisions and learn from experience. The advanced inspection system plans to control a robotic manipulator arm, an unmanned ground vehicle and cameras remotely, automatically and autonomously. There are many computer vision, image processing and machine learning techniques available as open source for using vision as a sensory feedback in decision-making and autonomous robotic movement. My responsibilities for the advanced inspection system are to create a software architecture that integrates and provides a framework for all the different subsystem components; identify open-source algorithms and techniques; and integrate robot hardware.
An integration of 3-D vision systems with robot manipulators will allow robots to operate in a poorly structured environment by visually locating targets and obstacles. However, by using computer vision for objects acquisition makes the problem of overall system calibration even more difficult. Indeed, in a CAD based manipulation a control architecture has to find an accurate mapping between the 3-D Euclidean work space and a robot configuration space (joint angles). If a stereo vision is involved, then one needs to map a pair of 2-D video images directly into the robot configuration space. Neural Network approach aside, a common solution to this problem is to calibrate vision and manipulator independently, and then tie them via common mapping into the task space. In other words, both vision and robot refer to some common Absolute Euclidean Coordinate Frame via their individual mappings. This approach has two major difficulties. First a vision system has to be calibrated over the total work space. And second, the absolute frame, which is usually quite arbitrary, has to be the same with a high degree of precision for both robot and vision subsystem calibrations. The use of computer vision to allow robust fine motion manipulation in a poorly structured world which is currently in progress is described along with the preliminary results and encountered problems.
Reflected computer vision targets are a powerful tool for measurement of mirror surface shape, with several important advantages over traditional fringe deflectometry methods. This method was first presented in 2021 and has undergone significant improvement and demonstration since. We describe a new baseline system using reflected computer vision targets, and present results from a large-scale measurement campaign conducted on both commercial heliostats and test mirrors in the laboratory. Calibration of the measurement system with photogrammetry allows for accurate measurement without careful control of target shape or camera position. Overall, the results show that a baseline setup using this method achieves measurement uncertainties in the slope error root-mean-square less than ±0.11 milliradian due to a series of repeatability conditions, varying sample position, rotation, lighting, camera settings, and system rebuild and recalibration. We present a detailed description of the setup, the results generated by this measurement tool, repeated measurement results, and the strengths and limitations of this metrology system.
Advanced Air Mobility (AAM) aircraft require precision approach and landing systems (PALS) in several environments, such as urban, suburban, and rural. It is challenging to implement current state-of-the-art methods approved for automated approach and landing for AAM operations with challenges such as GPS degradation in urban environments and visual navigation aids like the glideslope and localizer being narrow and not allowing alternative incoming landing angles at vertiports. However, existing technology and systems, i.e., the instrument landing system (ILS) with glideslope and localizer indicators that use vision, IR, radar, or GPS methods, provide baseline perception and sensing requirements for AAM aircraft approach and landing. This paper focuses on vision-based PAL and computer vision feature correspondence methods to demonstrate a baseline navigation system while adhering to the Federal Aviation Administration requirements and regulations about heliport design (FAA AC 150/5390-2C), which is one of the closest references for vertiport requirements and regulations. The coplanar pose from orthography and scaling with iterations (COPOSIT) algorithm determines pose estimation, which feeds into an Extended Kalman filter that combines IMU with vision to create a vision-based approach and landing (VAL) sensor fusion navigation solution for GPS-denied environments. The VAL navigation solution provides promising simulation results for AAM PALS with Hough circle detection and feature correspondence, which demonstrate robustness to false positives. This paper incorporates moderately high- fidelity simulations with computer graphics rendering to show a distributed sensor network to track an AAM aircraft during approach and landing to compare with the aircraft’s onboard vision-based navigation solution.
Computer vision is regarded as one of the most complex and computationally intensive problems. An integrated vision system (IVS) is considered to be a system that uses vision algorithms from all levels of processing for a high level application (such as object recognition). A model of computation is presented for parallel processing for an IVS. Using the model, desired features and capabilities of a parallel architecture suitable for IVSs are derived. Then a multiprocessor architecture (called NETRA) is presented. This architecture is highly flexible without the use of complex interconnection schemes. The topology of NETRA is recursively defined and hence is easily scalable from small to large systems. Homogeneity of NETRA permits fault tolerance and graceful degradation under faults. It is a recursively defined tree-type hierarchical architecture where each of the leaf nodes consists of a cluster of processors connected with a programmable crossbar with selective broadcast capability to provide for desired flexibility. A qualitative evaluation of NETRA is presented. Then general schemes are described to map parallel algorithms onto NETRA. Algorithms are classified according to their communication requirements for parallel processing. An extensive analysis of inter-cluster communication strategies in NETRA is presented, and parameters affecting performance of parallel algorithms when mapped on NETRA are discussed. Finally, a methodology to evaluate performance of algorithms on NETRA is described.
Integration of computer vision with on-board sensors to autonomously fly helicopters was researched. The key components developed were custom designed vision processing hardware and an indoor testbed. The custom designed hardware provided flexible integration of on-board sensors with real-time image processing resulting in a significant improvement in vision-based state estimation. The indoor testbed provided convenient calibrated experimentation in constructing real autonomous systems.
The response to the effects of nuclear detonations is supported by models that describe the evolution of the nuclear fireball and cloud and the associated transport of active debris. Validation of those descriptions relies on data from the nuclear test operations. Video records of those events offer a rich source of information that was exploited to a limited extent in historic analyses. Computer vision and machine learning techniques are powerful tools that can be used to increase the number of measurements that can be obtained from those films. In this work, we apply computer vision techniques to automatically track the temporal evolution of the nuclear fireball. In particular, we apply You Only Look Once 11 (YOLO11) and Segment Anything Model 2 (SAM2) in combination with minimal human intervention to digitized versions of the original nuclear test films. As part of the proposed workflow, the YOLO11 model is applied to films to determine bounding boxes for the fireball within each frame. These are then used as inputs to SAM2, which uses image segmentation to determine the fireball boundaries and their temporal evolution. We assess the accuracy of our approach by using it to determine the energy released during the Trinity nuclear test and comparing the results with previous analyses based on manual measurements.